Case study
Predicting how a hurricane cascades through US critical infrastructure
A hurricane does not damage one thing. It damages a port, which idles a refinery, which starves a region — and the asset records that would let anyone see that coming lived in dozens of federal and commercial datasets that shared no identifier.
- Client
- Cybersecurity and Infrastructure Security Agency (CISA), US Department of Homeland Security
- Sector
- Federal government · Homeland security · Critical infrastructure
- Role
- Prime contractor
- Period
- 2018 – present
- contractor role, not a sub
- Primecontractor role, not a sub
- the entity-matching engine built for it
- fastEMthe entity-matching engine built for it
- infrastructure sectors linked
- 5+infrastructure sectors linked
The problem
CISA is responsible for understanding what a disaster will do to the country's critical infrastructure before it does it. The blocker was not modelling. It was that the asset records needed to answer the question were spread across dozens of federal and commercial datasets that had been built separately, for different purposes, with no identifier in common.
Until those records are reconciled, there is no such thing as a forecast. You cannot predict what a storm does to a facility if the same facility appears four times under four names and you cannot tell that it is one facility.
What we built
The first half of the work is the half nobody photographs: ingesting, cleaning and normalising heterogeneous data — the IP Gateway, All Hazards Analysis, NOAA, US Census, Harvard Business School economic data, and sector-specific asset holdings — and then resolving it into one coherent picture.
- fastEM
- An entity-matching engine built for this, combining probabilistic linkage (fastLink), spatial reasoning (Haversine distance with K-D tree indexing) and deterministic rules to link and de-duplicate infrastructure records that share no key. Geography does work here that string comparison cannot: two records naming the same site differently are still at the same coordinates.
- Critical Asset List and Significant Asset Set
- Rule sets applied over the resolved records, and continuously refined, across banking, communications, energy, transportation and healthcare — delivered in standard formats the client's own analysts can work with rather than through a tool only we could operate.
- The forecast layer
- Models that project the impact of a natural disaster onto the assets identified, and onto the economic clusters that depend on them.
- The Infrastructure of Concerns Tool
- A secure cloud-hosted GIS platform: interactive maps, filtering by sector and region, reporting, and predictive-impact visualisation — plus the architecture documentation, function-level docs for fastEM and workflow guides that let it be maintained by somebody other than the people who wrote it.
Why the linkage is the whole engagement
Every model downstream inherits the quality of that reconciliation. If two records for one substation stay separate, the forecast counts it twice and overstates resilience. If two different substations get merged, it understates the exposure. Neither error announces itself in the output — which is why the matching engine, not the forecasting model, is where the effort went.
It is also why the same capability appears under a different name on almost every other engagement on this site. Record linkage across systems that were never designed to be joined is the most consistently underestimated part of applied analytics, and it is the part that decides whether the rest of it means anything.
Result
A cloud platform that lets decision makers see, before a storm makes landfall, which critical assets and which economic clusters are in its path. The client reports improved readiness and preparedness, better resource allocation through prioritisation and crew positioning, and reduced preparedness costs.
Those are the client's words rather than a measurement of ours, and they are stated without numbers because the numbers attached to this engagement are contract values and they are not ours to publish.
Tell us what decision you are trying to get right.
Not a discovery call about our capabilities. A conversation about the specific thing you need to predict, and whether the data you have can support it.