Case studies/Frontend Web Applications Platform

Public sector·Container platform, private cloud & storage·Romania

Frontend Web Applications Platform

A complete container platform, private cloud and storage layer for the secondary data centre — built as a peer to the primary site, with separate clusters for development, test and production, each split front-end and back-end.

ClientNational Tax Administration Agency
RoleArchitecture, integration & delivery
PlatformRed Hat OpenShift + OpenStack
StorageCeph + S3 object store, 240 TB
FootprintSecondary site
StatusDelivered
Frontend Web Applications Platform — secondary-site container platform, private cloud and storage

01The challenge

A second site that can carry the same platform

The estate’s container platform, private cloud and storage ran from a single primary data centre. Services used by citizens and businesses could not be left depending on one building, one power feed, one network path.

The mandate was to stand up the same platform in a second data centre — at the secondary data centre — capable of operating on its own or alongside the first, without changing the applications running on top of it.

P1A peer site, not a spareThe secondary data centre runs the same container platform, the same private cloud and the same storage as the primary — able to operate standalone, active-active or active-standby.
P2Container-firstA container platform on bare metal is the primary workload substrate; the private cloud sits underneath it for virtual machines and network functions, not the other way round.
P3Environments kept apartDevelopment, test and production run as separate clusters, each split front-end and back-end, so a change in one never reaches another by accident.
P4HA on every planePower, network, control plane and storage are each redundant inside the site, with segmentation policy between workloads.

02Scale

Sized for the platform,
built to lose components.

14
Nodes per site — 10 compute + 4 management
240 TB
Usable object storage
2
Peer data centres
40 Gbps
Uplink into the existing core

Per site: a four-node management cluster on bare metal, ten hyper-converged nodes running compute and storage together, a five-node object platform, and two redundant switch pairs dual-homed into the existing core.

03The platform

Eight layers, one platform

The container platform is the centre of this build. Everything else — the private cloud, the storage, the fabric — exists to serve the workloads running on it.

Management clusterOpenShift, four nodes, bare metal

Three combined control-plane and worker nodes plus one dedicated worker. The control-plane nodes carry etcd, the API server, the scheduler, the controller manager and the cluster version operator; the fourth node runs workloads only. Three control-plane nodes give etcd its quorum, so the cluster survives losing one.

Platform servicesRegistry, automation, fleet management

The management cluster hosts what the rest of the estate depends on: advanced cluster management across the fleet, an image registry, an automation platform, machine management, lifecycle management and a directory service — plus the software-defined networking, DNS and routing that the clusters themselves run on.

Workload clustersDevelopment, test and production

Application workloads run on their own clusters, separated by environment rather than by namespace alone, and each environment is split into a front-end cluster and a back-end cluster. Promotion between environments is a deliberate act; a back-end change cannot take a front-end with it.

Private cloudOpenStack, control plane as pods

The private-cloud control plane — identity, compute API and scheduler, networking API, images, orchestration, placement, dashboard, block and shared-file services, telemetry — runs as distributed pods on the management cluster rather than on its own controllers, so it inherits the same high availability and the same lifecycle.

Compute and storageTen hyper-converged nodes

Ten mixed nodes per site run compute and storage together: hypervisor instances, virtual networking agents and volume services alongside the distributed storage daemons. Capacity for virtual machines, network functions and persistent volumes grows by adding nodes.

Distributed storageCeph, through the storage operator

Block, file and object storage on the same nodes that run the workloads, presented to the container platform through its storage operator — so a pod asking for a persistent volume and a virtual machine asking for a disk are served by the same pool, replicated across nodes.

Object storeS3, for the long tail

A separate five-node object platform with a standard S3 interface holds long-term log retention and backup, with data replicated across all nodes so losing one node does not lose information. It is consumed primarily by the container platform for archival.

Network fabricRedundant switch pairs

Two switch pairs per site — one for 10/25G server connectivity, one for 40/100G — each configured as a virtual port-channel pair and dual-homed into the existing core network at 40 Gbps.

04Architecture

Management cluster on top, storage underneath

The diagram shows the reference architecture model — not the actual internal implementation. Each site runs its own management cluster, which provisions and governs the workload clusters below it; the private cloud and the storage layer sit underneath, and the two sites share one operating model with replication between them.

05Cluster model

Why the clusters are separate

Separation is the cheapest safety mechanism a platform has. Environments do not share a cluster, and inside each environment the front-end and back-end tiers do not either.

Management planeOne OpenShift cluster per site governs the fleet: it provisions the workload clusters, holds the registry and the automation, and runs the private-cloud control plane as pods.
DevelopmentIts own cluster. Build and integrate without a path into anything that serves the public.
TestIts own cluster, shaped like production, so what is verified is what ships.
ProductionIts own cluster, the only one carrying live workloads, on the same platform as the other two.
Front-end clustersPresentation and API-facing workloads, kept on separate clusters from the services behind them.
Back-end clustersBusiness services and data-facing workloads, reached only through the front-end tier and the platform’s own network policy.

06Storage

One storage layer, every access mode

Hyper-convergence means the nodes that run the workloads also hold the data. That removes a whole tier of hardware — and makes capacity a single planning exercise rather than two.

One pool, two consumersThe same distributed storage serves persistent volumes for containers and disks for virtual machines, so capacity is planned once rather than split between two estates.
Block, file and objectAll three access modes come from the same cluster, removing the need for a separate filer or a separate object system for day-to-day work.
Resilient by replicationData is replicated and distributed across nodes; losing a node costs capacity, not information.
Grows by adding nodesAdditional drives and additional nodes are in-place operations, with data redistributed automatically as capacity is added.
Archive on S3A dedicated five-node object platform takes long-term logs and backup off the primary tier, through a standard S3 interface.

07Resilience

What survives a failure

Two sites, one modelBoth data centres run the same platform stack with the same operating model; the secondary can run standalone, active-active or active-standby.
Replication between sitesThe storage layers replicate between the two data centres, so the secondary site holds the data as well as the platform.
Quorum inside the siteThree control-plane nodes per management cluster keep etcd’s quorum; the platform survives the loss of a node without operator action.
Segmentation between workloadsNorth-south traffic is inspected at the perimeter; east-west traffic between workloads is governed by segmentation policy inside the platform.
PerimeterA web application firewall and delivery tier sits in front of the estate — covered in its own case study rather than here.

08Confidentiality note

We do not disclose addressing, host naming, rack layouts or the actual internal technology architecture of the project.

The diagram and figures on this page describe the reference architecture model and the publicly communicable scale of the engagement.

09Next step

Standing up a second site?

We designed, integrated and delivered a container platform, private cloud and storage layer for the secondary data centre of a critical public estate. If you are planning geographic redundancy for a Kubernetes platform — a peer site rather than a spare — we can help.

Connect with us →