Skip to content
All insights

AI and Institutions

When AI Becomes Infrastructure

Author
Sarthak Joshi
Published
Reading time
19 min
Length
4,144 words
Embedded Systems Series · Paper 3 of 5

Systemic risk

When AI Becomes Infrastructure

Each AI system is judged by how often it fails. A society that runs on many of them should care more about how often they fail together, and about what its institutions have left to fall back on when they do.

In brief

  1. Systemic risk comes from correlation: shared providers, models, cloud regions and update channels. Making each system more reliable does not remove it. Framework
  2. Among 100 institutions that each fail 2% of the time, a pairwise correlation of only 0.01 makes a joint failure of a tenth of them about 440 times more likely. Model
  3. On one illustrative path, each institution's chance of failing falls roughly tenfold while the loss reached once a century rises about nineteenfold. Model
  4. AI governance checks systems one at a time. It needs a second, macroprudential layer that watches common exposures and fallback capacity. Framework

How claims are tagged: Evidence Model Inference Hypothesis Framework

One change, everywhere at once

On 19 July 2024 a faulty content update to a widely used endpoint-security product was live for 78 minutes. It affected about 8.5 million Windows devices, which Microsoft put at under 1% of all Windows machines. Hospitals cancelled surgery, an emergency-call system in Alaska was disrupted, and the Reserve Bank of India reported minor disruptions at ten banks and non-bank financial companies. Parametrix estimated the direct loss to the US Fortune 500, excluding Microsoft, at $5.4 billion, of which no more than 10 to 20% was insured.1 Evidence

The product was not an AI system. The event still shows the structure this article is about. One change went out through one channel to institutions in many sectors and countries. Each had chosen the same supplier for sound reasons. They failed together, although nothing else connected them.

AI is acquiring the same structure. Amazon, Microsoft and Google held 63% of the world's cloud infrastructure services market in the second quarter of 2026, and about 65% of India's cloud market in 2024 on the only commercial estimate available.2 Lower layers are more concentrated still. Two firms exclusively host the name servers for over 30% of the 10,000 most popular web domains, and five host about 60% of their index pages.3 Each AI deployment gets close attention for its accuracy, robustness and uptime. The correlation between deployments gets far less, and no measurement of it across deployed systems was found. Correlation is what makes a risk systemic.

Two different quantities

Individual risk belongs to one deployment: how likely it is to fail, and how much harm a failure does. Model evaluations, conformity assessments, service agreements and incident reports all measure it, and the remedy is a better system.

Systemic risk belongs to a population of deployments: how likely many are to fail at once, and the harm that simultaneous failure adds. It has three drivers. Common exposure means deployments that rely on the same provider, model, cloud region or update channel fail together. Coupling means one institution's failure spreads to those that depend on it, as when a payment processor stops and merchants stop with it. Depleted absorptive capacity means the manual processes, spare suppliers and trained staff that would have absorbed a failure have been retired. Framework

Figure 1

One sensible decision, two opposite effects

An institution moves from its own system to the market-leading modelA better model, better maintained, used by many others
Its individual risk fallsThe model is more accurate and more reliable than what it replaces.
Systemic risk risesIt now shares that model's failure modes and update schedule with everyone else who uses it.
Common exposureSame provider, model, cloud region, chips, data or update channel
Coupling and lost fallbackFailures spread along dependencies, and retired fallbacks cannot absorb them

Nobody records the second effect. The institution's own risk assessment improves, so the choice looks unambiguously good from inside.

Correlation sets a floor under system-wide risk

Take 100 institutions, each of which fails to deliver an AI-dependent function in a given year with probability 2%. Suppose each one's fate is copied, with probability w, from a single common shock that itself fails 2% of the time; otherwise it is drawn independently. Call w the common-shock loading. Every institution's own risk stays at 2% whatever w is. The pairwise correlation between any two institutions is w², because two institutions share the shock only when both copy it.4

Variance of the share failing = q(1 − q) × [ ρ + (1 − ρ) / N ]

As the number of institutions N grows, the second term vanishes and the first does not. Adding institutions, or adding redundancy inside each one that leaves the common exposure in place, removes only the independent part of the risk. The correlated part is a floor. Model

The tails move far more than the variance. Try the loading below.

Figure 2

Same individual risk, very different system

0 means fully independent failures. Each institution still fails 2% of the time.
0.01pairwise correlation (w²)
1 in 66chance that 10 or more fail in the same year
×440that chance compared with independent failures
1.8 × 10−8chance that 30 or more fail together
Independent failuresWith the chosen loading

Near the average of two failures a year the two curves are almost the same, which is why ordinary monitoring sees nothing. They part by many orders of magnitude in the tail, where a tenth or more of the institutions fail at once. Hover or tap the chart to read values.

Stated model, exact computation; 100 institutions, each failing with probability 0.02. Parameters are illustrative.

Individual risk held fixed, systemic risk varied100 institutions, each failing with probability 0.02
Loading wPairwise correlationSpread of the failing share (SD)P(10 or more fail)P(30 or more fail)
0 (independent)00.0143.4 × 10−58.0 × 10−27
0.050.00250.0163.1 × 10−36.9 × 10−14
0.100.010.0201.5 × 10−21.8 × 10−8
0.300.090.0442.0 × 10−21.3 × 10−2

Nothing about any single institution changes from row to row. An evaluation of each deployment finds the same 2% in every row, so every tool that assesses AI deployments one at a time is blind to the columns that change.

Market concentration is the wrong measure

The most visible source of correlation is concentrated supply. Suppose institutions spread across providers according to market shares, each provider suffers a disabling failure with probability 1% in a year, and each institution's total risk is again held at 2%. The implied pairwise correlation rises with the usual concentration index, the Herfindahl–Hirschman index: from about 0.05 when 100 institutions are spread evenly over ten providers to about 0.37 when one provider holds 85%.5

It would be natural to read the concentration index as a measure of systemic exposure. It is not, and the reason matters for regulation.

Figure 3

Which suppliers could push failures past the threshold on their own

Ten equal providers

  • Concentration index 0.10
  • P(30% or more fail) about 1 in 9,000
  • P(50% or more) 2.4 × 10−8

Four providers: 40 / 30 / 20 / 10

  • Concentration index 0.30
  • P(30% or more fail) about 2 in 100
  • P(50% or more) 4 in 10,000

Three providers: 60 / 25 / 15

  • Concentration index 0.45
  • P(30% or more fail) about 1 in 100
  • P(50% or more) about 1 in 100

One dominant: 85 / 15

  • Concentration index 0.75
  • P(30% or more fail) about 1 in 100
  • P(50% or more) about 1 in 100

Each column is one provider's market share. Red columns can take 30% of institutions down alone. The four-provider market has the lowest concentration index of the last three and the highest chance of a 30% failure, because two of its providers can each breach that threshold alone. At a 50% threshold the ranking reverses, because none of its providers can.

Stated model, exact enumeration over provider failures; 100 institutions, each failing with probability 0.02; each provider fails with probability 0.01.

What decides exposure at a given threshold is the set of suppliers whose failure alone would push the share of failed functions past the level at which failure becomes systemic, together with how likely each is to fail. A concentration index averages over all shares and cannot see a threshold. That gives a principled basis for designating systemically important AI suppliers: by whether their failure alone would breach a stated threshold of critical functions. It is a question about the dependency map, not about market share in the abstract. Model Framework

Providers are only the most visible shared element. Institutions that use different providers may still share a cloud region, a chip vendor, a base model fine-tuned by several vendors, a dataset or an update pipeline. Each shared element forms a failure group: a set of nominally separate dependencies that count as one when you assess redundancy.6

Figure 4

Six suppliers on paper, one failure group in practice

Institution 1
Institution 2
Institution 3
Institution 4
Institution 5
Institution 6
Institution 7
Institution 8
Supplier A
Supplier B
Supplier C
Supplier D
Supplier E
Supplier F
Cloud region Xused by A, B and D
Chip vendor Yused by C and D
Base model Zused by B, E and F

Procurement counts six suppliers. Because suppliers B and D each depend on two shared elements, the cloud region, chip vendor and base model chain all six suppliers into one failure group. The same grouping rule is used for subsea cables that share a landing corridor.

Schematic with generic labels.

The providers' own incident reports show the same structure inside their platforms. In October 2025 a latent race condition in the automation that manages DNS for one database service in Amazon's Northern Virginia region spread to compute, load balancing, serverless and sign-in services, in an event lasting about fourteen and a half hours. In June 2025 a single policy change with blank fields crashed a control service on which more than seventy Google Cloud products depended. In December 2024 a telemetry deployment overwhelmed the orchestration layer of OpenAI's clusters, on which their internal name resolution also depended; caching delayed the symptoms until the rollout had begun across the fleet.7 Each is a failure of a shared control plane, a component that everything built on it trusts, so one fault or one change reached all of it at once. Evidence

The reliability paradox

So far individual reliability was held fixed. In practice it improves, and what happens to systemic exposure depends on what institutions do with the improvement. Consider one illustrative ten-year path. Each institution's annual chance of a significant failure falls from 10% to about 1% as models mature. Adoption rises from under a tenth of the relevant work to over nine-tenths, so the fallback each institution keeps shrinks from 93% of the work to 7%. And as the market consolidates, the common-shock loading rises from 0.02 to 0.30.9 Better systems, used more widely, supplied by fewer firms, with less held in reserve.

Figure 5

Each institution gets safer while the system gets more fragile

Each institution's annual failure risk

Fallback kept and shared exposure

Fallback shareLoading w

Losses across 100 institutions

Expected annual lossLoss reached 1 year in 100

Losses are counted in institution-equivalents of unserved function. The expected loss rises from 0.66 to a peak of 1.76 in year six and ends at 1.0, about one and a half times where it began. The loss reached or exceeded one year in a hundred rises from 1.26 to 23.3, almost nineteen times.

Stated model on one illustrative path; it shows that expected and tail losses can diverge, not that they will.

Nothing in the expected-loss figure reveals the change, and that figure is what an institution weighs when it decides. Each choice along the path is sensible on its own terms: retiring a parallel process that is rarely used, consolidating on the best supplier, trimming staff whose skills are no longer exercised. Each lowers the institution's costs, and consolidation lowers its chance of failure. What none of these choices prices is the correlation it adds, and the tail that correlation creates is borne by the system as a whole. Flood defences that encourage building behind them work the same way: rarer floods, each doing more damage. Model

One outage, very different damage

Systemic risk is about many institutions failing together, but the damage lands institution by institution, and it varies widely. The programme's national outage loss model shows why. It treats each sector's daily loss during a connectivity outage as output, times the share of it that needs connectivity, times the share the manual fallback cannot cover, times the share never recovered. The fallback is a stock that runs down: cash in the till, fuel in the generator, paper forms, the supervisor who remembers the old process. In the model's central case, finance keeps about a fifth of its first-day fallback by day five and agriculture about two-thirds. A national week without connectivity costs India an estimated $17.1 billion in direct losses (range $9.6 to 26.0 billion), and the seventh day costs 1.34 times the first.10 Model

The model was built for connectivity, and it describes sectors rather than institutions. Its structure can still be borrowed, because three of its four parameters are exactly what AI adoption changes: how much of a function depends on the system, how much the fallback can carry at first, and how long the fallback lasts. Inference

Figure 6

Two institutions, the same disruption

A seven-day outage: share of the function lost each day

Warm fallback: covers 60%, runs down slowlyThin fallback: covers 20%, runs down fast

A 72-hour outage: share of the function delivered

70% within 2 hours, 6-hour recovery15% after 2 days, 1-day recovery

Left: with the same dependence on the failed service, the warm fallback loses 2.5 function-days over the week and the thin one 3.9. Right: the institution with a warm fallback loses about 24 value-hours, the other about 79, 3.3 times as much. Raising the second institution's fallback to 70% closes just over a third of the gap; matching the first's switchover time closes about an eighth, and its recovery time about a seventh. The parameters interact.

Stylised comparisons with illustrative parameters.11

The outage record gives these parameters content.12 Evidence

  • Independence is what counts. At Matsu in 2023 a microwave link, independent of the cut cables, carried voice and state communications for about fifty days. At Tonga in 2022 the eruption that cut the cable also blocked the satellite fallback. A 2026 review of the largest North American grid-security exercise found backup communications "often underpinned by the same technology" as the primary.
  • Dependence travels. During the West African cable cuts of 2024, cloud collaboration services were disrupted even in countries not directly hit.
  • Domestic switches are common modes too. On 12 April 2025 India's unified payments interface ran at about 50% success for two hours and about 80% for three more; one report attributed it to banks' status-check requests reaching the switch without a rate limit.
  • Existing fallbacks are small. UPI Lite, the small-value wallet within which India's offline payments sit, carried 0.38% of UPI transactions in December 2024. How much is genuinely network-free is not published.
  • Consequences go unmeasured. The official inquiries into India's 2012 grid failure, the 2003 North American blackout and the 2025 Iberian blackout contain no consequence data.

How badly a common failure hurts an institution depends largely on the independence, capacity, engagement time and endurance of its fallbacks, and these are rarely measured while things work. The most direct way to know them is to test them: periodic, announced operation without the AI dependency, measuring how much of the function continues and for how long. Framework

Objections, briefly

"Markets will price the risk."

They will to the extent that the institution that chooses bears the risk. A correlated failure falls on everyone exposed, while the saving from consolidation goes to each institution separately. That is the structure of an externality, and macroprudential regulation exists because it did not correct itself in finance.

"Diversity costs more than it buys."

For most functions, yes, which is why the proposals below are confined to critical ones. For those, the expected-loss calculation that makes diversity look expensive is the calculation that cannot see the tail.

"AI providers are more reliable than what they replace."

Often true, and the reason adoption is rational. Provider reliability and system resilience are different quantities, and improving the first can weaken the second when institutions shed reserve and converge on one supplier.

"This argues against sovereignty."

The companion paper recommends that a middle power hold the switches critical functions pass through. A held switch is still a single point: India's April 2025 payment outage was a domestic failure. The answer is to hold the switch and make it redundant, with independent instances and offline modes for the most critical transactions.

A macroprudential layer for AI

AI governance today works one system at a time. It evaluates models, certifies systems, audits deployments and logs incidents. The argument here is that it also needs a layer concerned with the population of deployments. Financial regulators have already built part of the vocabulary. The EU's supervisory authorities designated critical ICT third-party providers for direct oversight in November 2025, and the UK's regime for critical third parties took effect in January 2025.13 Their criteria, such as systemic impact, concentration of reliance and substitutability, carry over to AI suppliers with little change.

Common-exposure register

For critical functions, record which providers, models, cloud regions, chip vendors and update channels each institution relies on, and resolve them into failure groups. No institution can build this map alone.

Designation by breach capacity

Designate as systemically important any supplier whose failure alone would push affected critical functions past a stated threshold, and supervise it accordingly.

Staggered updates

Release model updates in stages across institutions, and let critical functions hold back until an update has run elsewhere. A simultaneous failure becomes a sequential one that can be stopped.

Fallback standards and drills

Require critical institutions to state and test the independence, capacity, switchover time and endurance of their fallbacks, and to report the results like other resilience measures.

Shared regression baselines

Run a common set of test cases against the models that critical institutions use, before and after each update, so that silent, correlated changes become visible.

For India

Pieces exist: protected payment systems, six-hour incident reporting, telecom critical-infrastructure rules, and 109 cyber-security mock drills with 1,438 organisations by March 2025.14 None is yet organised around common exposure to AI suppliers.

What this analysis cannot tell you

Every model here is stated, not estimated, and its parameters are illustrative. The single-factor structure is the simplest that shows the effect; real dependence is layered, with institutions sharing a cloud region inside one provider and a base model across several. The reliability-paradox path is one of many, chosen to show that expected and tail losses can diverge. The outage model describes sectors, not institutions, and has never been tested against a national event, because none has happened. Above all, the most important quantity identified here, the actual degree of correlation among AI deployments in any country, has not been measured. The article also leaves deliberate attack aside, although it exploits the same common exposures.

The question usually asked about AI reliability, how often a system will fail, is the right question about one system and the wrong one about a society that runs on many systems built on the same foundations. For a society, the questions are how failures are correlated, which suppliers could take down a critical share of functions on their own, and what each institution keeps to carry on when they do.

Notes

  1. CrowdStrike, remediation and guidance hub (update live 04:09 to 05:27 UTC); D. Weston, Microsoft, 20 July 2024, for the 8.5 million devices and the share of Windows machines; Parametrix, 24 July 2024, for the loss estimate. The cancelled surgeries, the Alaska emergency-call disruption and the Reserve Bank of India statement come from the programme's verified incident record. ↩
  2. Synergy Research Group, 30 July 2026 (28, 20 and 15%). For India, the Competition Commission of India's market study on AI (October 2025), citing Market Research Future: AWS 32.6, Azure 20.8 and Google Cloud 11.5%. It is a single commercial estimate. ↩
  3. Wang et al., measurement of the Tranco top 10,000 domains (2021, revised 2024), consistent across six vantage points. ↩
  4. A one-factor model of exchangeable failures. Tail probabilities are computed exactly by conditioning on the common shock. Parameters are illustrative, and the model establishes structure rather than magnitude. ↩
  5. Provider failures are enumerated exactly for 100 institutions; each provider fails with probability 0.01 and each institution's total risk is held at 0.02. ↩
  6. The grouping rule comes from the companion study's cable exposure index, which treats cables sharing a landing corridor, a submarine canyon or a duct as one group. It is a proposed instrument and has not been back-tested. ↩
  7. Amazon Web Services, summary of the DynamoDB service disruption in the Northern Virginia (US-EAST-1) Region, 19 to 20 October 2025; Google Cloud incident report, 12 June 2025; OpenAI incident report, 11 December 2024. ↩
  8. Koh et al., WILDS benchmark, Table 1 (Camelyon17, average accuracy in and out of distribution). ↩
  9. The fallback share follows the steady state of the reversibility relation in the first paper of the series with its decay and rebuilding rates set equal; neither rate has been measured for any deployment. The loading of 0.02 to 0.30 is a pairwise correlation of 0.0004 to 0.09. The "loss reached once a century" is the number of failures reached or exceeded with probability at most 0.01, times each failed institution's unserved share. ↩
  10. The companion study's national outage loss model; its parameters are its authors' judgements. The day-seven to day-one ratio is 1.42, 1.34 and 1.25 in its low, central and high cases. ↩
  11. Seven-day case: dependent share 0.6; fallback covering 60% with a characteristic time of seven days, against 20% with two and a half days. Seventy-two-hour case: a switchover-and-recovery form without depletion. ↩
  12. The programme's incident register and facts log: Matsu (2023) and Tonga (2022); NERC and E-ISAC, GridEx VIII Lessons Learned Report (March 2026); West Africa (2024), from the Internet Society, ThousandEyes and Bloomberg. The UPI outage is verified; its success rates and cause rest on a single report. UPI Lite shares are derived from the Reserve Bank of India's Payment System Report (December 2024) and NPCI statistics. The programme's record carries the Matsu duration with an unresolved conflict against later repair dates. ↩
  13. EIOPA, 18 November 2025; Bank of England, PRA and FCA, PS16/24 (UK regime in force from 1 January 2025). The number of EU designations comes from a secondary source. ↩
  14. Press Information Bureau, 28 March 2025. The drill count is cumulative, with no start date given. Protected-system designations (June 2022), telecom critical-infrastructure rules (November 2024) and CERT-In reporting directions (April 2022) come from the programme's facts log. ↩
Selected sources
  • Acemoglu, D., Ozdaglar, A. and Tahbaz-Salehi, A. (2015). Systemic risk and stability in financial networks. American Economic Review.
  • Bommasani, R. et al. (2021). On the opportunities and risks of foundation models. arXiv:2108.07258.
  • Buldyrev, S. V. et al. (2010). Catastrophic cascade of failures in interdependent networks. Nature 464, 1025–1028.
  • Cobbe, J., Veale, M. and Singh, J. (2023). Understanding accountability in algorithmic supply chains. FAccT 2023, 1186–1197.
  • Kleinberg, J. and Raghavan, M. (2021). Algorithmic monoculture and social welfare. PNAS 118(22).
  • Koh, P. W. et al. (2020). WILDS: a benchmark of in-the-wild distribution shifts. arXiv:2012.07421.
  • Tessone, C. J. et al. (2013). How big is too big? Critical shocks for systemic failure cascades. Journal of Statistical Physics 151, 765–783.
  • Wang, S. et al. (2021; revised 2024). Measuring the consolidation of DNS and web hosting providers. arXiv:2110.15345.
  • Laha, A. et al. (2026). India's Internet Sovereignty. Working paper P01.

Adapted for general readers from Working Paper P17 of the Embedded Systems Series (September 2026). The working paper carries the full argument, the provenance of every figure and the complete reference list. Model results are computed from the paper's stated models and are not estimates of any real system.