Back to blog

IRIS — Integrated Risk Intelligence

By Rplus AnalyticsCase study12 Aug 2026

The fraud and error service Rplus built on the Data Services Platform for a central government department, identifying £340M of organised fraud.

£340MOrganised fraud identified
72 → 3Hours to run the match
45 TBMatched under risk rules
38+Datasets in discovery

What it is

IRIS, the Integrated Risk Intelligence Service, is one of three tenants on the Data Services Platform. Its purpose is to identify fraud and error across the whole department, spanning its various benefits.

Rplus took IRIS through discovery, alpha and beta, and delivered both the IRIS risk engine and the address matching MLOps pipeline.

What went into discovery

Data discovery for IRIS covered more than 38 datasets — drawn not only from the department itself, but from other government departments, local councils and national reference data.

Datasets identified in discovery
  • Maternity allowance
  • Disability benefits
  • Winter fuel benefits
  • Council datasets
  • PAYE data — another government department
  • Immigration data — a third department
  • National postcode data
Blue: departmental data. Amber: sourced from outside the department — which is where matching gets hard, because nothing shares a key.

Three days became three hours

The service required address matching across the country, drawing from multiple sources, running over 45 terabytes of data under multiple complex risk rules.

In a normal instance that query took three days to run.

TIME TO RUN THE NATIONAL ADDRESS MATCH Before a normal instance 72 hrs After on the platform 3 hrs The blue block is the whole of the new runtime, drawn to the same scale.
Same 45 TB, same complex risk rules. The query was developed as Hive SQL on top of the data, carrying those rules with low latency.

After moving to the Data Services Platform, it was reduced to three hours.

How the matching works

Python provides the address matching functionality that retrieves fraud and error across the various benefits. The whole matching process runs as an MLOps pipeline.

SOURCES 38+ datasets 45 TB RISK RULES Hive SQL low latency MATCHING Python address matching CLUSTERING R · K-Means similar fraud groups Risk engine the whole matching process runs as an MLOps pipeline
Rules, matching and clustering are separate stages — which is what makes the process something that can be run, monitored and rerun rather than executed by hand.

Data science

R is used extensively for the service. Clustering algorithms developed in R, including K-Means, identify groups of citizens making a similar type of fraud — the pattern that distinguishes organised fraud from isolated error.

Unstructured data from Universal Credit is extracted through MongoDB to perform the risk intelligence work.

Worth noting

The department is one of the largest users of MongoDB in Europe. More than 20 data scientists work on the project, with a Principal Data Architect leading a team of over 20 data analysts, scientists and engineers.

Hosting and security

Because the service handles very sensitive data covering citizens across the country, hosting was determined by risk profile. Based on high risk profiling of the data, IRIS was hosted on premise with full security controls in place.

The result

£340M of organised fraud identified in a financial year
Proof points
  • £340M of organised fraud identified
  • 72 hours reduced to 3 hours
  • 45 TB matched under complex risk rules
  • 38+ datasets in discovery
  • 20+ data scientists