Back to blog

The Data Services Platform

By Rplus AnalyticsCase study11 Aug 2026

A central government department's enterprise data lake, built by Rplus from the discovery phase in 2017 and still supported today.

2017Built from discovery
240+Source systems
200+ TBHeld on the platform
3Tenants, five services

What it is

The Data Services Platform is a big data platform serving divisions across the department's digital service as a multi-tenancy platform. Three tenants run on it: IRIS, the Integrated Risk Intelligence Service; ChADS, Children Analytics Data Services; and GySP, delivered as Data as a Service. Route2Cloud and Apply for NINO are also delivered on the platform.

Rplus provided architecture and development as a service to build DSP from scratch. We have been part of the project since the discovery phase in 2017, working with stakeholders to develop the long-term roadmap and through delivery into sprint planning. IRIS, ChADS, GySP and Route2Cloud were each taken through discovery, alpha and beta. Alongside the platform itself we delivered the IRIS risk engine and the address matching MLOps pipeline.

The scale of it

DSP holds more than 200 terabytes. It draws on more than 240 source systems, with over 250 datasets from internal and external departments mapped into the platform and landed into a mix of SQL and NoSQL databases.

Alongside DSP we worked extensively on the department's central data warehouse — its corporate memory, holding 25 years of history.

SOURCES Internal systems Other departments Call centres Legacy warehouses 240+ systems 250+ datasets INGEST Informatica BDM Apache NiFi Spark SQL · Hive SQL MS SSIS PROCESS Hadoop Spark Spark Streaming STORE SQL · NoSQL Hive · HBase Oracle · PostgreSQL 200+ TB IRIS  ·  ChADS  ·  GySP plus Route2Cloud and Apply for NINO Where volumes were too large for batch, Spark Streaming handles ingestion into Hive and HBase.
One platform, many sources, several services — the reason segregation between tenants had to be designed in rather than added later.

The engineering

We designed and delivered the data engineering patterns that extract data from those source systems, and implemented Informatica DEI from scratch. CI/CD data pipelines cover end-to-end design, development, test and deployment.

Integration spans Cloudera, Oracle, PostgreSQL, flat files, R and Airflow. MS SSIS extracts data from the department's call centres into the lake. Postgres supports Django applications running on the platform. Pipeline design uses technical information at row and table level to support GDPR obligations, and we implemented a business glossary and data catalogue covering the department's programmes.

Ingestion
Informatica BDMInformatica DEIApache NiFiSpark SQLHive SQLMS SSIS
Processing
HadoopSparkSpark StreamingDatabricks
Storage
HiveHBaseOraclePostgreSQLNoSQLCloudera
Orchestration
AirflowAzure Data FactoryInformatica DEI
Serving
Qlik SenseDjangoR
Delivery
GitHubCI/CDInfrastructure as codeTwo-week sprints

The platform follows a two-week sprint and release model with a fully automated DevOps CI/CD pipeline. Code is managed through GitHub, with merge requests reviewed and approved by tech leads before promotion. More than three development teams from different suppliers work on the same codebase.

Keeping tenants apart

Because divisions share one platform, segregation between onboarded tenants was a core requirement. A breach between tenants would carry major repercussions.

It was achieved by agreeing resource allocation per tenant across every dimension where they could otherwise interfere with one another, with strict access controls using RBAC and ABAC.

ONE SHARED PLATFORM IRIS ChADS GySP ALLOCATED SEPARATELY PER TENANT Physical hardware Virtual machines Software Batch windows Maintenance windows RBAC and ABAC access control
Segregation is not a single control. It is an agreement about five things, enforced together.
A constraint worth naming

Budget and resources made building the full solution in one go impossible. Tenants were onboarded one at a time, following departmental standards, until all were live.

Hosting and migration

DSP is hosted on the department's own data centres and on AWS and Azure. Over five years of migration work we moved data from legacy systems and obsolete data warehouses onto the platform, including a large-scale programme that transferred 260 TB between the department's two data centres. Personal identifiable information moving from on-premise to AWS was handled under encryption and obfuscation policies.

We also delivered infrastructure as code to spin up the development environments, fully automated through the CI/CD pipeline, with environments shut down and restarted according to usage without data loss. Twenty large development and test clusters were optimised and virtual machines right-sized to reduce monthly cost.

Security and governance

National Insurance numbers, tax reference numbers, names, gender and addresses are handled with encryption, obfuscation and masking. Security standards are considered at the earliest design stage. We worked with the department's cyber resilience teams on penetration testing aligned to the NIST Cyber Security Framework, and built the alerting and monitoring tooling for its Audit Trail Analysis Service and Cyber Resilience Centre.

NOTHING STARTS UNTIL IT IS APPROVED High and lowlevel designs Design Authorityapproval before work starts Build and implementgoverned through delivery Security assessedand approved Alongside: Digital Decision Authority for new patterns  ·  Technical Design Oversight Groups Delivery follows Government Digital Service standards, with governance taken as the starting point and worked back from.
High and low-level designs go to the Design Authority before work commences. Once approved, we govern those solutions through implementation and maintenance.

DSP was assessed and approved by the department's Enterprise Security Risk Management team against its security standards and policies. We work with department teams across networks, infrastructure, security, core services and identity, and alongside external suppliers including Deloitte and Agile Solutions, hardware suppliers Nutanix and HP, and software suppliers Cloudera, Informatica and Qlik. Twenty security-cleared Rplus staff work on the platform.

Resilience

We delivered the disaster recovery solution for DSP, working with the department's technology services, database administration, networking, servers and DevOps teams, to enable data recovery after a disaster.

Running it now

Rplus provides post-go-live warranty support for DSP. Delivered statements of work include:

Disaster recovery solution design
Maintaining the Qlik Sense visualisation layer
Building and maintaining the Informatica ELT layer
Maintaining the Cloudera big data layer
Maintaining the platform's infrastructure as code

Each with knowledge transfer and training for departmental staff.

Proof points
  • Built from discovery in 2017
  • 240+ source systems
  • 250+ datasets mapped
  • 200+ TB held
  • 260 TB migrated between data centres
  • 25 years of history in the central data warehouse
  • 3 tenants, 5 services
  • 20 security-cleared staff