Case studies/APIA — Infrastructure rebuild

Public sector · Compute, storage & virtualisation · Romania

APIA — Infrastructure Rebuild

Storage that had fallen out of support, carrying data that could not wait. 34 servers, two all-flash arrays and a redundant network across two sites — with the live data migrated off the failing estate underneath running workloads.

ClientAPIA · Public sector, Romania
RoleSupply, installation & configuration
PlatformRed Hat Enterprise Virtualization
SubstrateDell EMC PowerEdge + Unity all-flash
FootprintTwo sites, 5 cabinets
StatusDelivered
Rebuilt agency data-centre infrastructure — all-flash storage arrays and dual-homed servers across two sites

01The challenge

Replacing the floor while people are standing on it

APIA’s central IT node was running on storage that was ageing out of vendor coverage. A considerable part of the existing capacity was no longer under warranty or support, and the workloads on it were the agency’s own — they could not simply be paused.

The task was to install a new platform across two sites, move the live data onto it with as little interruption as possible, and leave behind something that extends by adding hardware rather than by being replaced again.

P1 Storage falling out of support A significant share of the existing storage was no longer covered by warranty or vendor support — and the data on it could not wait for a procurement cycle.
P2 Migration without an outage window Moving live data off the failing array meant working across the storage layer and the hypervisor at the same time, keeping interruption to a minimum.
P3 Headroom, not a like-for-like swap Whatever replaced it had to be extensible later without another forklift — more capacity by adding shelves, not by starting again.
P4 Two sites, one operating model A second site in a state data centre had to run the same platform as the headquarters, managed the same way.

02Scale

What was installed,
across two data centres.

34
Servers across two sites
1,632
Physical cores
86 TB
Usable all-flash capacity
17 TB
Total memory

Five 42U cabinets, thirty-four dual-socket servers, two all-flash arrays holding sixty-two enterprise SSDs between them, four data-centre switches, two out-of-band management switches and a three-phase parallel-redundant UPS — delivered, racked, cabled and configured across a headquarters site and a secondary data centre.

03What was delivered

Seven layers, one delivery

Compute, storage, the data services on top of it, network, virtualisation and power were supplied and configured as one integrated solution — every licence included, every system at its current release on the day it went in.

Compute 34 dual-socket servers, all-SSD

Twenty-two hosts at the central site and twelve at the secondary, each with two 24-core processors, 512 GB of memory and five enterprise SSDs. Racked in pairs for airflow, across five 42U cabinets with dual power distribution.

Storage All-flash arrays, two redundant controllers each

One all-flash array per site — 40 SSDs giving 52 TB usable at the central site, 22 SSDs giving 27 TB at the secondary, both in RAID 5. Each array carries two hot-swap controllers with 128 GB of cache apiece, battery-backed so cache is flushed to flash on power loss, and every software feature licensed and enabled.

Data services Tiering, efficiency and point-in-time recovery

Inline compression and deduplication for both block and file volumes, thin provisioning, automatic tiering that moves hot blocks onto the fastest media, snapshots and thin clones for production copies, and journaled replication that allows recovery to any point in time, grouped per application so interdependent systems come back consistently.

Access Block and file, on the same array

Fibre-channel and iSCSI for block access alongside SMB, NFS, FTP and SFTP for file — so the array serves both the virtualisation platform and file workloads without a second system.

Network Redundant switch pair per site

Two data-centre switches per site with 10/25 Gbps access ports and 100 Gbps uplinks, configured as a virtual port-channel pair. Every device lands one port in each switch: a single switch can fail without dropping a path, and while both are up their capacity adds together — 20 Gbps of data per host. Storage traffic is isolated on its own VLAN that exists only locally.

Virtualisation Red Hat Enterprise Virtualization, 34 hosts

An enterprise Linux hypervisor on every server, all licensed for virtual datacentre use. Each site is managed by its own manager, which itself runs as a highly-available virtual machine on the cluster — if its host fails it restarts elsewhere, so the management plane survives the loss of a node.

Power Redundant, three-phase

A parallel-redundant 40 kVA three-phase UPS at the central site, with two power distribution units per cabinet, each fed from an independent UPS or generator source, so no equipment depends on a single supply path.

04Architecture

Every device dual-homed

The diagram shows the reference architecture model — not the actual internal implementation. The principle is visible in it: at each site, every device lands one port in each of two switches, storage is reached over four independent paths, and the management network is entirely separate from the data path.

Host connectivityThe first two interfaces on each server are bonded for data — capacity added, redundancy gained. Two further interfaces carry storage traffic, each with its own address, so a lost link does not cost the path to the array.
SwitchingA virtual port-channel pair per site; every device is dual-homed with one port into each switch. Uplink to the institution’s existing network is a redundant port-channel with each fibre landing in a different module.
Storage pathFour 10 Gbps iSCSI ports per array — two per controller — split across the two switches, so both a controller and a switch can be lost without taking storage offline.
SegmentationStorage traffic runs on a dedicated VLAN that exists only on the local switches and is never carried beyond them.
ManagementA separate out-of-band switch reaches the management interfaces of the hosts, the switch pair and the array, so the platform stays reachable even when the data path is not.
Between the sitesA single carrier-provided 10 Gbps layer-2 circuit — documented as the one link in the design without a second leg, landing in one switch at each end.

05Migration

Moving data off a failing array

The migration was the difficult part of this project, not the installation.

Two layers at onceMigrating live data off failing storage means understanding both the array and the hypervisor: volumes are moved underneath running machines rather than by taking the machines down.
Non-disruptive by designThe replacement array takes firmware and software updates without interrupting service, so maintenance after go-live does not reopen the same problem.
Native to the hypervisorThe array integrates with the virtualisation platform directly — capacity is provisioned from the management console, virtual machines are visible against the volumes they sit on, and copy and move operations are offloaded to the array instead of running through the hypervisor.
Room to growCapacity extends by adding drives, with data redistributed automatically across the array — the reason the original estate had to be replaced does not recur at the next capacity ceiling.

06Outcome

What the agency ended up with

Data off the risk

The workloads that were sitting on unsupported, out-of-warranty storage now run on two fully-licensed all-flash arrays with redundant controllers and battery-backed cache.

A path that survives failure

Dual controllers, dual switches, dual power feeds, bonded host interfaces and dual storage ports — every layer of the delivered platform has a second leg.

Performance where it is needed

All-flash media with automatic tiering, inline compression and deduplication, and 20 Gbps of data per host.

An estate that extends

Additional drives, additional hosts and additional shelves are all in-place operations, so the next growth step is not another rebuild.

07Confidentiality note

We do not disclose site addressing, host naming, credentials, rack layouts or the actual internal technology architecture of the project.

The diagram and figures on this page describe the reference architecture model and the publicly communicable scale of the engagement.

08Next step

Storage running out of support?

We rebuilt the compute, storage and network platform of a national agency across two sites, and moved its live data off an unsupported array with minimal interruption. If you are carrying an estate that is ageing out of coverage, we can help.

Connect with us →