Skip to content
All insights

AI and Institutions

The Algorithmic State

Author
Sarthak Joshi
Published
Reading time
19 min
Length
4,209 words
Embedded Systems Series · Paper 5 of 5

Decision authority

The Algorithmic State

A system labelled "decision support" can end up making the decision without anyone choosing that it should. Override rates cannot tell you whether it has. A small randomised sample of cases can.

In brief

  1. Formal authority, who is entitled to decide, can stay with officials while de facto authority, what actually determines the outcome, passes to the system. Framework
  2. The test is whether overrides land on the system's errors. Three reviewers who all override 5% of cases can end with 5.0%, 8.2% or 12.2% error against the system's 8%. Model
  3. If unexercised judgement decays, an institution that stops practising looks fine until the system fails, then its error on recommended cases nearly doubles. Hypothesis
  4. One instrument answers all three problems: decide a small random share of cases without the recommendation. Framework

How claims are tagged: Evidence Model Inference Hypothesis Framework

Twenty thousand a week

When the Australian government automated the recovery of welfare overpayments in 2016, debt notices rose, on figures reported in the case literature, from about 20,000 a year to 20,000 a week.1 The system was designed to remove staff from determining and communicating debts; notices went out without a government employee in the loop. On the Royal Commission's figures it raised debts against 433,000 people, and the government reimbursed 746 million Australian dollars and wrote off claimed debts of 1.751 billion. The Commission called the scheme "a crude and cruel mechanism, neither fair nor legal". Legal advice before it began had said the method would not be supported by law.2 Evidence

Robodebt is often told as a story about a bad algorithm, and it was one. It is also a story about authority. Nobody decided each of the debts raised against those 433,000 people; the system did. Human judgement survived only upstream, in the design and the decision to deploy, and downstream, in the courts that undid it and the inquiry that examined it.

A government deploying AI today might say its systems are different: a recommendation is made, and an official decides. This article asks what that description secures. Its answer is that the passage of authority from official to system can be measured, though no measurement of it in deployed public review was found; that it can be individually rational and institutionally corrosive; and that agents that act, rather than recommend, can remove even the formal point at which a human decides.

A spectrum that hides the question

The algorithmic state is usually described as a spectrum of increasing automation, from a human decision, through AI-assisted and AI-recommended decisions, to automated decisions, agentic decisions and autonomous institutional action. The stages are defined by who is entitled to decide. Consequences depend on what determines the decision. The two can come apart at every stage from the second onwards, and most of all at the third, the one described as decision support. Framework

Figure 1

Formal and de facto authority along the spectrum

More automation
Formal authoritywho is entitled to decide
De facto authoritywhere it can migrate
Oversight must showto keep the two together
Human decision
Formal authorityOfficial
De facto authorityStays with the official
Oversight must showNothing extra: no system is involved
AI-assisted
Formal authorityOfficial
De facto authorityTo how the information supplied is framed and selected
Oversight must showThat the official's judgement shaped the result
AI recommendation"Decision support"
Formal authorityOfficial
De facto authorityTo the recommendation, when it is accepted without discrimination
Oversight must showThat overrides track the recommendation's errors
Automated decision
Formal authorityOfficials, through rules the system applies
De facto authorityTo designers, vendors and whoever sets thresholds
Oversight must showThat rules, thresholds and exceptions are chosen and answerable
Agentic decision
Formal authorityOfficials, through a goal and bounds delegated to the system
De facto authorityTo the goal specification and the agent's intermediate choices
Oversight must showThat the actions stay within stated bounds and can be reversed
Autonomous institutional action
Formal authorityUnclear
De facto authorityTo interacting systems across agencies
Oversight must showThat some authority answers for the combined effect

Formal authority (grey) and de facto authority (amber) coincide only at the first stage. The highlighted stage is the one usually called decision support: the recommendation is formally advice, and in practice may be the decision.

After Table 1 of the working paper. Framework

What the literature already knows is mixed. Public administration spotted the shift early: Bovens and Zouridis described street-level bureaucracies becoming system-level ones, with discretion moving from front-line officials to system designers.3 The UK's data-protection regulator has said a decision does not escape the rules on solely automated processing merely because a human rubber-stamped it, and the EU AI Act names automation bias as the one cognitive bias overseers must be enabled to guard against.4 The behavioural evidence cuts both ways. Dutch public-sector experiments with 605 and 904 citizens and 1,345 civil servants found no evidence of automation bias: most participants overrode an algorithmic score when other evidence contradicted it. They did find selective adherence: a Moroccan-Dutch teacher with a low score was, in the authors' words, 50% more likely not to have a contract renewed than a Dutch teacher with the same score.5 Green's review of 41 human-oversight policies argues both that people struggle to perform the oversight required and that the policies lend legitimacy to faulty systems.6 The first claim is overstated, since people sometimes do oversee well. The second stands. The problem is not that officials never exercise judgement. It is that nothing in the formal arrangement shows whether they do. Evidence

Making de facto authority measurable

Suppose a recommender errs on a share e of cases. An official overrides a wrong recommendation with probability h and a right one with probability a. The official's oversight discrimination is D = h − a: how far overrides concentrate on the recommender's errors. At D = 0 the official holds formal authority and has delegated it de facto, whatever the override rate, while adding errors of their own. Estimating D requires knowing, for a sample of cases, whether the recommendation was right. Framework

Final error = e(1 − h) + (1 − e)a  ·  Override rate = e h + (1 − e)a

Figure 2

The same override rate, three kinds of authority

The system errs on 8% of cases, and every official shown overrides 5% of them.
1.1%right answers overridden, a
0.49oversight discrimination, D
5.0%final error, against the system's 8%
Adds valueverdict on this official

All points on the line override exactly 5% of cases. The first marked official catches half the system's errors and rarely overrides a right answer (5.0% final error). The second exercises some judgement (D = 0.27) but does slightly worse than the system alone (8.2%). The third overrides at random with respect to the system's errors (D = 0) and adds errors while appearing to supervise (12.2%).

Stated model with illustrative parameters. Model

Deference can be rational

A reviewer improves on a recommender only if they catch more errors than they introduce. When the recommender is accurate, that is demanding. An official who catches half the errors of a system that is wrong 5% of the time improves on it only if they override fewer than 2.6% of its correct recommendations; if the system is wrong 2% of the time, the ceiling is 1%. An official above the ceiling does better by deferring, and experience will tell them so: in the cases where they overrode, the system will more often than not turn out to have been right. What looks like automation bias may be a correct response to base rates. That is the first reason de facto delegation is hard to prevent. Model Inference

Figure 3

How often an official may overrule a right answer and still add value

Maximum share of the system's correct recommendations an official can override and still improve on it, for an official who catches half of the system's errors. The more accurate the system, the less room for independent judgement to pay: 1.0% at a 2% error rate, 2.6% at 5%, 5.6% at 10% and 16.7% at 25%.

Stated model: improvement requires h e > a(1 − e). Model

Delegation by drift

The difficulty lies in what rational deference may do over time. The hypothesis is that judgement not exercised decays: an official who has deferred for years has less independent skill than when they started, and discrimination depends on skill. If so, the damage is hard to see while the recommender performs well and decisive when it does not. Hypothesis

A stated model makes this concrete. A reviewer's skill, between zero and one, sets both the willingness and the ability to override, so a reviewer who loses skill defers more rather than overriding at random. Skill is rebuilt by deciding cases without the recommendation and lost through disuse. Then the recommender's environment shifts: its error rate is 5% for five years and then rises to 25%, because the population changed, a policy it encodes changed, or the model was updated.8

Figure 4

Skill kept by practice, lost by disuse

Rate at which skill is lost, relative to how fast practice restores it
0.56skill after five years (full skill = 1)
4.4%error on recommended cases before the shift
15.0%error on recommended cases after it
3.9% → 12.6%override rate, before and after

Reviewer skill over five years

No practice10% practice30% practiceYour setting

Error on recommended cases, before and after the shift

BeforeAfter

With no practice, skill falls from 0.9 to 0.27 over five years; with a tenth of cases decided independently it holds at 0.56, and with three-tenths at 0.81. Before the shift, the error on recommended cases is 4.7%, 4.4% and 4.1%, so the institution that stopped practising looks almost as good, and cheaper. After it, the errors are 20.2%, 15.0% and 10.6%. Override rates, the signal that something has gone wrong, rise about threefold everywhere: from 1.9% to 6.0% with no practice, and from 5.6% to 18.1% with three-tenths.

Stated model with illustrative parameters. It counts error only on cases decided with the recommendation; the independently decided cases carry a cost the model does not count.9 Model

The evidence for the mechanism is suggestive, not conclusive. In Allegheny County, less experienced child-welfare call screeners aligned more closely with an algorithm's scores than senior ones, though the ability to compensate for miscalculated scores was not associated with experience.10 A scoping review of 937 papers describes an "efficiency–atrophy paradox" while stressing that the evidence for AI specifically points only to a hypothesised erosion.11 No study found measures skill decay among officials using AI, or reports override rates in deployed public-sector review. Hypothesis

The capacity paradox

AI can raise the capacity of the state. In a November 2023 survey, British NHS professionals thought generative AI, properly exploited, could cut the share of their time spent on bureaucracy from 50 to 30%; algorithmic targeting of homelessness prevention in Los Angeles County was, on a preliminary assessment, associated with substantial reductions in homelessness.12 These are perceptions and early results. No study reviewed measures a productivity gain from AI in public administration.

Two qualifications apply. Some gains are transfers: the governance series' paper on triage proves that in a queue that never idles while work waits, prioritising cases cannot reduce average delay, only move it; at 90% utilisation, prioritising the top fifth cuts their wait to 0.12 of first-come-first-served and lengthens everyone else's by a fifth.13 And capacity to act and capacity to err scale together. Michigan switched on an automated unemployment-fraud system with a single administrative decision; reversing its effects on the roughly 40,000 people it falsely accused took nine years, legislation, litigation and a $20 million settlement.14

Accuracy makes deference rational, speed makes independent practice costly, and both erode the human capacity that would catch the system's failure.

The state's capacity rises in normal conditions and grows brittle in abnormal ones. It is the institutional analogue of the reliability paradox in the infrastructure paper of this series. A British survey experiment found the public side of the same pattern: efficiency gains initially raised trust in AI-assisted government while reducing citizens' sense of control.15 Inference

When the decision point disappears

Agentic systems change the problem in kind. A recommendation has a decision point, the moment a human accepts or overrides it, where discrimination can in principle be measured. An agent pursuing a goal through a sequence of actions may have no such point, or many at which review is nominal. The evidence is conceptual and legal rather than empirical: interviews with civil servants suggest agents intensify oversight problems built around siloed compliance units and episodic approvals; an analysis of more than 1,300 benchmark papers found none that meets public-sector requirements; and legal analyses find the AI Act fits agents poorly.16 No study reports outcomes for agentic AI deployed in government. For agents, oversight has to move from decisions to trajectories: the bounds within which an agent may act, the actions it may not take without confirmation, and the ability to reverse a sequence once begun.

Six cases

Where authority went, and who corrected it
CaseWhat was automatedHow the failure was corrected
Robodebt, AustraliaIncome averaging inferred overpayments, and the burden of proof moved to recipients; in what appears to have been about 80% of cases they could not produce acceptable evidence.A class action and a Royal Commission, years after the harm began.
Childcare benefits, NetherlandsA risk model used nationality, and dual nationality specifically, as a risk indicator; over 25,000 people were wrongly accused on one count. The data-protection authority fined the tax authority a record €2.75 million.A regulator and political inquiry, years later.
SyRI, NetherlandsA welfare-fraud risk-profiling system.On 5 February 2020 the District Court of The Hague held its legal basis contrary to Article 8(2) of the European Convention on Human Rights, on privacy grounds rather than on the quality of any decision.
A-level grades, UK 2020With examinations cancelled, an algorithm downgraded 39% of teacher-assessed grades using schools' past results.After public outcry, the regulator reverted on 17 August 2020 to the higher of the teacher-assessed and moderated grades.
MiDAS, MichiganRecipients could be designated fraudulent automatically; about 40,000 people were falsely accused.Litigation, legislation and a settlement, nine years later.
Food distribution, India98.2% of electronic ration transactions in July 2026 were biometrically authenticated, for 79.2 crore (792 million) beneficiaries. In August 2016 in Ranchi district, only 52% of ration-card households managed to buy their ration.A formal fallback for network failure exists; no measurement of its use was found, and no correction is recorded.

In none of the first five cases was the failure corrected by the oversight built into the process. Correction came from outside the decision system: courts, inquiries, regulators and public outcry.17 In India, where an authentication failure is itself a decision about who eats, no correction is recorded at all. Evidence

One instrument, three jobs

The central recommendation is a randomised independent-decision sample. For each class of decisions made with an AI recommendation, a small share of cases is assigned at random to officials who decide without seeing the recommendation. Framework

Figure 5

How the randomised sample works

Cases in one decision classfor example, reviews of benefit eligibility
A small share, chosen at randomdecided by officials who do not see the recommendation, which is still recorded
All other casesdecided with the recommendation, as now
Outcomes of the sampled casesobserved or independently adjudicated, and set against the recommendations and the officials' decisions on recommended cases
1 · Keeps judgement in practicethe fallback when the system fails
2 · Measures discrimination, Dshows whether overrides track the system's errors
3 · Supplies the audit holdoutoutcomes for cases the system would have screened out

The auditing paper in the governance series shows that when outcomes are observed only for cases a system let through, no sample size repairs the gap; a randomised holdout is the only general fix, and it "requires an authorisation nobody currently seeks".18 On the drift model, a tenth of cases sustains about half of full skill if skill decays a tenth as fast as practice restores it, but less than a fifth if it decays half as fast. Its cost in staff time, and in any extra error on the sampled cases, is bounded by the share chosen; it is not estimated here.

What else to do

Report discrimination, not override rates

For each decision class: the system's error rate, the override rate, and the share of overrides that corrected an error. With a well-specified record the override rate is a simple query; the other two need each case's outcome, or its later reversal, joined to the record.

End the fiction where measurement shows it

Where discrimination is near zero, either reinvest in oversight or reclassify the decisions as automated, with the contestation rights, explanation duties and legal review that automated decisions attract.

Use calibrated probabilities

An overconfident forecast used at a nominal threshold can do worse than acting on the base rate alone; recalibration takes days and restores a small positive value. A reviewer inherits a recommender's miscalibration.

Size audits to their question

Detecting a five-point gap in false-positive rates for a group that is 5% of cases takes 18,283 cases at conventional power, and testing twenty subgroups with error control roughly halves the power of a fixed sample.

Bound agents by trajectories

Specify what an agent may do without confirmation, what needs it, and how a sequence of actions is reversed, and budget reversibility as the first paper in the series proposes.

Keep a route to challenge

Houston teachers whose evaluations rested on an undisclosed method had to go to court to win on due process. The AI Act's new right to explanation has contours one analysis calls "under-defined".19

Logging for accountability has a privacy cost too. Statistics published from such logs under a formal privacy guarantee spend a lifetime budget that is spread thinner the longer a programme runs and the more often it publishes; on the governance series' privacy paper, a 30-year annual programme leaves every cell under 843 people as noise at 5% relative error.20 Retention limits and access controls should be designed in from the start.

Objections, briefly

"Officials are not rubber stamps."

The Dutch and Allegheny evidence shows people overriding algorithmic advice, mostly when other evidence contradicted it. De facto delegation is not inevitable. It cannot be assumed away either, because nothing in the formal arrangement reveals it, and the experimental nulls come from vignettes, not caseloads.

"Humans override badly, so involve them less."

Overrides can inject bias, as in Kentucky's bail decisions, and a review of laboratory studies finds people often override algorithms for the worse.21 The case for independent practice does not rest on humans being better on average. It rests on keeping a capacity that can catch the system's failures, and on the measurement only independent decisions allow. Where judgement is worse, that argues for a small sample rather than human review of every case.

"Keep the human for legitimacy."

An experiment in Finland found that transparency and human discretion raise the perceived legitimacy of automated decisions, more among administrators than among citizens.22 That is the mechanism Green warns about: a human in the loop supplies legitimacy whether or not they supply judgement. Legitimacy resting on authority not exercised in fact lasts only until the gap is exposed.

What this analysis cannot tell you

The models are stated, with parameters never estimated for any official or system. The drift model counts error only on cases decided with the recommendation, so the cost of the practice it recommends is not measured, and measuring discrimination needs outcomes that many administrative decisions censor. The behavioural evidence comes mostly from experiments and one well-studied deployment. The cases come from high-income democracies and India, and several of their figures differ between sources.

"Decision support" describes formal authority. It says who is entitled to decide and nothing about who decides. The remedy proposed here is modest: decide some cases without the machine, at random, and use them to keep judgement alive, to see whether authority is exercised, and to learn what the system's errors are.

Notes

  1. Clarke, Michael and Abbas (2024), citing media reporting; 169,000 notices had been sent by January 2017. ↩
  2. Clarke, Michael and Abbas (2024), quoting the Royal Commission's report: 433,000 individuals; $746 million reimbursed to about 381,000 people; $1.751 billion written off. The class-action settlement (over $751 million repaid) and the Prime Minister's figure of more than 500,000 victims are different quantities and are kept separate. Legal advice of November 2014, as recorded by the Commission. ↩
  3. Bovens and Zouridis (2002), abstract. ↩
  4. Green (2022), quoting the UK Information Commissioner's Office (2020) and the Article 29 Working Party (2018); Laux and Ruschemeier (2025) on Article 14(4)(b) of the AI Act, whose text was read through scholarship. ↩
  5. Alon-Barkat and Busuioc (2022): 12.3 against 8.6%; odds ratio 1.50, one-sided p = .04, whether the advice came from the algorithm or a human expert; not found among civil servants surveyed after the childcare-benefits scandal. ↩
  6. Green (2022), Computer Law & Security Review. ↩
  7. The governance series' working paper The Accountability Gap in Automated Administration (2026), sections 3 to 5; results on stated models. ↩
  8. Skill changes as dk/dt = ηu(1 − k) − κ(1 − u)k, with learning rate 0.20 and decay rate 0.02 a month, initial skill 0.9, and overrides scaling with skill (maximum catch rate 0.8, false-alarm scale 0.03). All parameters are illustrative. ↩
  9. The model implies that override rates fall as skill decays: with no independent practice, from about 6.2 to 1.9% over five years. That is a prediction of the hypothesis, not evidence that such a decline happens; no study found reports override rates in deployed public-sector review. ↩
  10. Cheng and Chouldechova (2022), shown and corrected scores from December 2016 to July 2018. ↩
  11. Žnidarič, Štrumbelj and Machidon (2026). ↩
  12. Bright et al. (2024), a survey its authors caution is not fully representative; Dedyaev (2026), preprint, no effect size given. ↩
  13. The governance series' working paper Algorithmic Triage in Public Service Delivery (2026): a stated queueing model with one pooled server. ↩
  14. Dedyaev (2026), preprint; Stewart (2026) on the 2024 settlement. ↩
  15. Wuttke, Rauchfleisch and Jungherr (2025): a preregistered experiment with 1,201 British respondents. ↩
  16. Schmitz, Rystrøm and Batzner (2025); Rystrøm et al. (2026); Gardhouse, Oueslati and Kolt (2026); Reed et al. (2026). ↩
  17. Fraser and Stardust (2025) on litigation in the SyRI, Michigan and Robodebt cases; Pinsent Masons (2021) on the Dutch fine; Rechtbank Den Haag, ECLI:NL:RBDHA:2020:1878; Butler (2025) for the 39% figure, not independently verified; Ofqual statement of 17 August 2020; IMPDS and Press Information Bureau for the Indian figures; Menon (2017), abstract, for Ranchi. ↩
  18. The governance series' working paper Auditing Public-Sector Models (2026), negative result 6.2 and the holdout discussion. ↩
  19. Lyons, Velloso and Miller (2021); Pavlidis and Kastanas (2025). The calibration and audit-size results come from the governance series' papers on calibration and on auditing, on stated models. ↩
  20. The governance series' working paper Privacy Budgets for Public Analytics (2026), Table 4: a lifetime budget of ε = 1 over 30 annual releases. ↩
  21. Albright (2019), as reported in Cheng and Chouldechova (2022) and Green (2022); Green (2022). ↩
  22. Hillo, Vento and Erkkilä (2025): 842 Finnish administrators and 3,245 citizens. ↩

Adapted for general readers from Working Paper P19 of the Embedded Systems Series (September 2026). The working paper carries the full argument, the provenance of every figure and the complete reference list. Model results are computed from stated models and are not estimates of any real institution.