AI and Institutions
Beyond the First Order
- Author
- Sarthak Joshi
- Published
- Reading time
- 22 min
- Length
- 4,632 words
Higher-order consequences
Beyond the First Order
We measure what AI does to a task. Institutions are changed by what happens next: who gets hired, what knowledge gets written down, what capacity survives. Here is a way to reason about those consequences, and a test of how far that reasoning can be trusted.
In brief
- A capability changes how activities connect, not how much happens. Its effect is exposure times leverage, so the same tool can help one office and hurt another. Model
- When capabilities change as fast as institutions respond, analysis goes stale. At plausible settings, reasoning stays trustworthy for fewer than two steps. Model
- Under that churn, better analysis buys little. Redoing analysis more often probably buys more. Inference
- Adoption wears away the capacity to reverse it, so reversibility has to be budgeted like money. Framework
How claims are tagged: Evidence Model Inference Hypothesis Framework
The first order is the part we can see
Most of what has been measured about generative AI is first-order. A customer-support assistant raised issues resolved per hour by 14% on average and by 34% for novice and low-skilled agents. In a preregistered experiment, professional writing tasks took 40% less time and output quality rose 18%. Consultants using GPT-4 completed 12.2% more tasks, 25.1% faster, on work within the model's capabilities, but were 19 percentage points less likely to reach the right answer on a task beyond them.1 In a late-2023 survey of 938 British public-service professionals, 22% already used such tools and 32% said there was clear guidance on their use.2 Evidence
This evidence says little about what happens next. After ChatGPT's release, incumbent freelancers on a large online labour platform submitted 61.8% fewer bids and moved toward other work.3 Beyond that, the questions multiply. Does the employer of a more productive support agent still hire juniors? Is the knowledge those juniors would have gained still produced anywhere? Are public answer forums still being written? Can a ministry that reorganises its caseload around a drafting tool still work without it? These are the consequences governments face, and on them the evidence is thin.
Two habits fill the gap. Enumeration lists risks and benefits without the mechanisms that would produce them, so nobody can check it. Extrapolation projects a first-order number forward and says nothing about the reorganisations that decide whether the number matters. This article offers a method in between: connect a specified capability to institutional consequences through mechanisms that are written down, and say how far the reasoning can be trusted.
Order is a property of the description
"Consider the second-order effects" assumes that order is a fact about the world. It is a fact about a description. An effect is of order n if the shortest route from the intervention to it, through a stated set of mechanisms, has n steps. Merge two nodes of the description and the order falls; split one and it rises; nothing in the world has changed. The programme's companion paper on nth-order thinking proves this in general, and draws the practical rule: an order claim made without a declared set of mechanisms cannot be checked, and can be manufactured by choosing the level of detail.4
AI's consequences are usually told as a chain: capability, adoption, behaviour, organisation, institution, economy and society, geopolitics, and policy response. Read as a list of orders, that chain misleads, since an institutional effect can arrive in two steps or in six. Read as a list of layers, it is useful. The layers differ in how fast their mechanisms move, who drives them, and how visible they are to anyone evaluating a decision. Framework
The consequence chain, its delays, and the loops that close it
Loops that feed back
Delays broadly lengthen down the chain. Visibility is highest at the top and at enacted policy, and lowest in the middle, where behaviour, organisational capacity and institutional norms change. On this framework, much of the institutional consequence travels through that unmeasured middle.
Conceptual. R marks reinforcing loops and B balancing loops; the full inventory is below. Hypothesis
A capability changes connections, not quantities
A conventional policy instrument changes the level of something: a tax changes a price, a grant changes a budget. An AI capability does something different. It lowers the cost or raises the quality of an activity, and so changes how strongly one part of a system responds to another. In the language of the companion paper, it changes the mechanisms rather than the inputs. Standard sensitivity analysis then gives its effect on an outcome as a product of two factors.5
Effect of a capability ≈ leverage of what the mechanism feeds × change in the mechanism × exposure, the activity already flowing through it
Exposure is not consequence. Task-exposure studies measure the first factor. One estimates that about 80% of the US workforce could have at least a tenth of their tasks affected by large language models, and about 19% at least half.6 Their predictive power is contested: tested against US unemployment-insurance records for 2010 to 2020, before language models were in use, individual exposure indices did not predict unemployment or job separations, though an ensemble of them did.7 Exposure measures leave out leverage: how much a change in that activity moves the outcome once every downstream route is counted. A tool touching a large, low-leverage activity can matter less than one touching a small activity that feeds a consequential decision, such as drafting eligibility decisions or triaging the cases that use an institution's scarcest staff. Inference
The institution, not only the technology, sets the sign. Exposure records what is already flowing through an institution's mechanisms, so the same change to the same mechanism can raise or lower the outcome depending on what the institution is already doing. The companion paper's stated model of an office working to a performance target shows this without modification.
An office working to a target, and three tools that could enter it
Arrows carry the model's stated weights. Three tools act on three different links. Tool A makes measured work more useful to quality. Tool B absorbs the paperwork, weakening the way measured work crowds out unmeasured work. Tool C makes unmeasured work more productive.
The companion paper's stated five-mechanism target model, unchanged. Quantities are changes caused by the target.
Which link a tool strengthens decides whether it helps
Marginal effect on the target's net welfare effect of strengthening each link by one unit. Tool A helps most (+1.20) and Tool B also helps (+0.66). Tool C, the one a naive reading expects to help most, hurts (−0.74): the target has already pulled effort away from unmeasured work, and a stronger link from that work to quality makes the withdrawal costlier.
Stated model; not an observation. Values checked against numerical derivatives for every link. Model
A capability amplifies the flows it touches. Whether amplification helps depends on whether those flows currently run with or against the outcome. That is why first-order evaluations are informative about the mechanism and nearly silent about the consequence, and why the same tool produces different results in different organisations. Inference The practical rule does not depend on the linear model: name the mechanisms a capability changes, weight each by the activity flowing through it and the leverage of what it feeds, and treat the institution's current state as part of the answer.
How far ahead can you see a moving system?
The companion paper derives a ceiling on how deep analysis can usefully go. A route of n steps passes through n mechanisms that someone had to specify. If each is specified correctly with probability p, the whole route is right with probability pn. Requiring a route to be at least as likely right as wrong gives about three steps at p = 0.8, which the companion paper considers generous for policy work, and six and a half at p = 0.9.8 On that account, depth is bought by better specification.
That assumes the mechanisms stay fixed while effects travel along them. With AI they do not. A mechanism specified correctly today, say that a drafting tool saves officials time and the backlog absorbs it, can be invalidated by the next capability change before the effect has passed through it, for instance when the tool starts filing as well as drafting. If each mechanism takes some time to operate and each specification stays true for some time, the chance that every step of an n-step route is still right when the effect reaches it has a simple form.9
Route fidelity Φn = pn ÷ [(1 + ν)(1 + 2ν) … (1 + nν)], where ν = time a mechanism takes to operate ÷ time its specification stays true
Later steps are exposed to change for longer, so the penalty grows faster than exponentially with depth. Explore it below.
Usable depth when the system keeps changing
Chance a route of each length is still right
Usable depth as churn rises
Left: bars above the dashed line are at least as likely right as wrong. Right: the three curves converge as churn rises, so a better analyst buys less and less depth.
Stated model with illustrative parameters; checked against a simulation of delays and obsolescence at three points (error below 0.001). Neither p nor ν is a measurement.
| Chance each mechanism is right, p | ν = 0 | 0.05 | 0.10 | 0.20 | 0.50 |
|---|---|---|---|---|---|
| 0.70 | 1.9 | 1.6 | 1.5 | 1.2 | 0.9 |
| 0.80 | 3.1 | 2.3 | 1.9 | 1.5 | 1.1 |
| 0.90 | 6.6 | 3.3 | 2.6 | 1.9 | 1.3 |
Three things follow. The ceiling falls quickly. At the reference fidelity of 0.8, a mechanism that takes a twentieth of its useful life to operate cuts usable depth from 3.1 steps to 2.3, and one that takes a tenth cuts it to 1.9. Depth falls below two steps once ν passes 0.088, and a second-order route, capability changes behaviour and behaviour changes the organisation, is then already less likely right than wrong. Model
Better analysis stops paying. With fixed mechanisms, raising p from 0.8 to 0.9 buys three and a half extra steps. Under churn it buys one step at ν = 0.05 and about two-thirds of a step at ν = 0.10. What plausibly replaces it is cadence: analysis redone at intervals short relative to how long a specification stays true should face a smaller effective ν, because each refresh resets the clock on every mechanism it re-examines. The model itself contains no refresh, so this part is inference. Model Inference
Nobody has measured ν. It is plausibly small for mechanical links, since a data centre's energy draw does not change sign with a model release, and plausibly large for organisational and institutional links, whose operating times of quarters to years match the interval between major capability changes. That is testable: map a deployment's mechanisms, map them again after each major capability change, and count the links whose existence or sign changed. If ν proves large, commissioning one thorough assessment of an AI deployment is the wrong habit. Hypothesis
Adoption consumes reversibility
When no analysis can settle the sign of a decision's consequences, the companion paper's advice is to buy reversibility: staged rollout, time limits, a review date set after the decisive effects should arrive, and monitoring on the routes the analysis could not resolve.10 Fast-moving capabilities will need that advice more often. The difficulty is that AI adoption is unusually good at destroying what the advice relies on.
Reversibility is the retained capacity to perform a function without the capability: staff who still know the work, processes that still run, data held outside the vendor's system, suppliers who could still be engaged. Each is kept up by use and decays with disuse, and successful adoption is what stops the use. The parallel process is retired because it costs money, and knowledge of its edge cases leaves with the people who held it. Inference In a simple model where capacity decays with the share of work done by the capability and is rebuilt by the share still done without it, retained capacity depends only on that share and on the ratio of decay to rebuilding.11
Retained capacity falls with the share of work handed over
Retained capacity depends on how much work still passes through the alternative, not on how good the capability is, and it heads to zero as that share does. Reversing a deployment at high adoption means rebuilding from a low base while the function is done badly or not at all.
A framework, not an estimate: neither rate has been measured for any AI deployment. Framework
If reversibility is the remedy for decisions that cannot be analysed, and adoption consumes it, it has to be budgeted as deliberately as money: a stated floor of retained capacity for each critical function, a share of work routed through the alternative to keep it warm, and tripwires at which further adoption pauses until the floor is restored. This restates the control half of the Collingridge dilemma, that control is lost once a technology is entrenched, as a design parameter: entrenchment becomes something to measure and its rate something to choose. The infrastructure and algorithmic-state papers in this series develop two instances, fallback capacity in an outage and an official's ability to tell when a recommendation is wrong.
The loops that recur
The chain closes into loops, and a few recur across domains. Naming them gives analysts a checklist when they declare a mechanism map, and gives scenario work a principled basis: the future of a domain depends on which loops dominate. Evidence exists for some; others are stated without it.
| Loop | How it works | Status of the evidence |
|---|---|---|
| R1 Capability–investment flywheel | Adoption brings revenue, data and demonstrations that fund more capability. | Stated without evidence |
| R2 Complementarity lock-in | Adopters invest in processes, integrations and skills that raise switching costs. | Stated without evidence |
| R3 Fallback erosion | Retired alternatives stop being usable (see Figure 5). | Inference |
| R4 Symmetric automation | When one side of an exchange automates applications or complaints, volumes rise and the other side automates its screening. | Stated without evidence |
| R5 Access expansion | Cheaper expertise reaches people previously unserved. Whether AI compresses skill gaps is unsettled: novices gained most in customer support and less creative writers gained more from AI ideas, but a forecasting experiment found no consistent pattern.12 | Mixed evidence |
| B1 Incident-driven restriction | Visible failures bring restriction, but only failures at the visible ends of the chain. | Stated without evidence |
| B2 Commons depletion | Stack Overflow posting fell about 25% against comparable platforms within six months of ChatGPT's release, a lower bound; model collapse is its technical form.13 | Evidence, one platform |
| B3 Price and wage adjustment | Cheaper tasks lower prices and wages and redirect effort. | Stated without evidence |
| B4 Policy's own loop | A review of 41 policies requiring human oversight of government algorithms found people cannot perform the oversight prescribed, and argued the policies lend legitimacy to flawed systems.14 | Argued in a review |
| B5 Strategic response | Capability diffused to rivals prompts countermeasures and more investment. | Stated without evidence |
Two further structures belong in any map. Homogenisation: when many actors use the same few models, their outputs and errors correlate. A meta-analysis of 28 studies found people working with generative AI slightly better on creative tasks while the diversity of their ideas fell sharply, and a formal analysis shows that convergence on one algorithm can lower social welfare even when it is more accurate for each user.15 Performativity: a prediction that informs decisions changes what it predicts. In a model of predictive policing on Oakland data, patrols that learn only from incidents they themselves discover converge on sending all patrols to one of two areas, where matching true crime rates would send 56.7%.16
Building scenarios from loops rather than stories
Each loop leaves observable signatures. Rising volumes with falling acceptance rates signal symmetric automation. Falling contributions to public knowledge signal commons depletion. Falling entry-level hiring in exposed occupations signals erosion of the route by which expertise is produced. This is more modest than a forecast and more useful than a narrative.
Framework
A worked case: generative AI in entry-level work
Customer support, drafting, analysis and programming are where first-order evidence is strongest. Applied to them, the framework produces an order ledger that records, for each route, its sign, when it should arrive, the evidence behind it and what would settle it. Every route is written out step by step, so its order can be checked.
The evidence thins almost immediately after the first step
1 step
2 steps
3 steps
4 steps
Twelve routes from the abbreviated ledger, placed by the number of steps from the capability. Several unresolved routes are benign, and access expansion may dominate where expertise is scarce.
The entry-level route shows how unstable early evidence can be. US payroll data first showed a 13% relative decline in employment of 22-to-25-year-olds in the most AI-exposed occupations (August 2025), then 16% (November 2025); the current version (August 2026) puts their employment 19% below where it would have been had it kept pace with less exposed peers, while noting that the original regression result is no longer statistically significant, that the pattern weakens once education is controlled for, and that the facts are descriptive.17 Danish register data on about 25,000 workers in exposed occupations found precise null effects on earnings and recorded hours two years after ChatGPT's launch, ruling out effects above 2%, alongside real task reorganisation.18 That speaks to the wider labour market rather than to entry-level hiring itself. The route is contested at the level of the evidence, which is what a ledger should record. Evidence
Pilots are typically short. Singapore's government coding pilot ran from October 2023 to January 2024, and it reported coding time down 22% while senior developers warned that newcomers leaning on an assistant might compound mistakes.20 A pilot that ends before organisations adapt is informative about shallow effects and silent about deep ones. The unseen deep effects run both ways, with fallback erosion and commons depletion on one side and complementary investment on the other. An evaluation should therefore state its silence about the deeper routes in the evaluation itself. Inference
Objections, briefly
"Institutions adapted to electricity and computing without any of this."
They did, over decades; the first industrial dynamo appeared in the 1870s and its full effect on productivity in the 1920s. If organisations adapt as slowly this time, consequences arrive slowly enough to manage. But slow organisations do not make ν small. ν is operating time divided by specification life, so slow mechanisms raise it unless the capability they respond to also settles. The evidence on pace cuts both ways: individual use spread fast, to 39% of US adults aged 18 to 64 by August 2024, while measured effects on earnings stay small, and one task-based projection puts AI's ten-year contribution to productivity at no more than 0.66%.21
"Just adapt as you go."
The staleness result supports adaptive management. Adapting as you go requires the capacity to change course, however, and adoption consumes it. It is a property a deployment has to be designed to keep.
"Linear models cannot represent thresholds."
Correct, and the most serious limit here. A labour market that absorbs displacement until it does not cannot be captured by a constant weight. The framework can flag candidate tipping points, where loop gains approach one, but cannot locate them.
"The framing hunts for harms."
The calculation itself is sign-neutral, as Figure 3 shows, and the loop inventory includes benign loops. Benefit routes should be recorded with the same care as cost routes.
What to change
Specify capabilities as mechanism changes
Start an assessment by naming the links a capability alters, the activity flowing through each and the leverage of what it feeds. An assessment that describes accuracy and stops has described the link, not the consequence.
Give every assessment an expiry date
State the resolution, the ceiling on depth and whether the sign is settled, plus the interval after which the assessment must be redone, tied to the capability's release cycle rather than the budget cycle.
Budget reversibility
For functions an institution cannot stop, set a floor of retained non-AI capacity, route a share of work through the alternative, test it, and write exit terms into procurement.
Monitor the middle of the chain
Track entry-level hiring in exposed occupations, contributions to public knowledge, the share of work still done without AI, and the concentration of the models public functions rely on.
Extend attribution
Keep a retrievable record of each deployment decision: the mechanisms assumed, the effects expected, the routes left unresolved. It lets a distant harm be traced to an earlier claim.
Record benefits as carefully as costs
Access expansion and complementary investment are real routes. A ledger that only lists harms is as unreliable as one that only lists gains.
What this analysis cannot tell you
The apparatus is linear and local, so it cannot represent thresholds, saturation or interaction. Both computed results follow from stated assumptions: the staleness ceiling assumes independent delays and obsolescence with common rates, and correlated obsolescence, where one capability change invalidates many mechanisms at once, would make things worse. The target-office illustration reuses a stated model and shows a structure, not a magnitude. The parameter that matters most, ν, has not been measured anywhere. Most evidence on higher-order effects post-dates 2023 and comes from single studies whose replication is not yet known.
"What are the second-order effects of AI?" is a malformed question. Without a declared set of mechanisms it has no answer; with one, it has many that depend on the level of detail chosen. The answerable questions are narrower. Which links does a capability change, and how much flows through them? Which routes carry the consequence, and which does the evidence reach? How long does an analysis of them stay valid? And how much capacity to change course remains after adoption?
Notes
- Brynjolfsson, Li and Raymond (2023), NBER working paper; the published version (2025) reports 15% on average. Noy and Zhang (2023), Science. Dell'Acqua et al. (2023); the published version (2026) states the outside-frontier effect as 19% less likely. ↩
- Bright et al. (2024): an online survey in November 2023, which its authors caution is not fully representative. ↩
- Yiu et al. (2024, version of May 2025): 312,143 incumbent freelancers on one platform; 51.2% in the specification adjusted for pre-trends. The authors note the sample selects highly active freelancers. ↩
- The programme's working paper Nth-Order Thinking for Policy and Government Decisions, section 2.3. ↩
- Standard adjoint sensitivity analysis on the companion paper's linear form: the first-order change in the outcome from a change in the link from j to i is the leverage of i times the change times the activity at j. ↩
- Eloundou et al. (2023), arXiv version abstract; published in Science (2024) as "GPTs are GPTs: labor market impact potential of LLMs". ↩
- Frank, Ahn and Moro (2023), abstract. ↩
- Companion paper, sections 1.1 and 4.1. ↩
- Mechanism delays and obsolescence times are independent and exponentially distributed; the expectation factorises into the product shown. Non-integer depths use its continuous extension. ↩
- Companion paper, sections 4.4, 4.5, 7.3 and 12. ↩
- Steady state of dR/dt = −σaR + η(1 − a)(1 − R), giving R* = 1 / [1 + (σ/η) · a/(1 − a)]. ↩
- Brynjolfsson, Li and Raymond (2023; 2025); Doshi and Hauser (2024), Science Advances; Ng et al. (2024) for the government coding pilot (juniors 33% against an average of 22); Schoenegger et al. (2024). ↩
- del Rio-Chanona, Laurentsyeva and Wachs (2024), PNAS Nexus; Gao et al. (2026) on how 40 platforms govern AI-generated content; Shumailov et al. (2024), Nature, on model collapse. Whether production training pipelines are exposed to the same degree is not established. ↩
- Green (2022), Computer Law & Security Review. ↩
- Holzner, Maier and Feuerriegel (2025): Hedges' g = 0.27 for performance and −0.86 for diversity, the latter on six observations with high heterogeneity. Kleinberg and Raghavan (2021), PNAS. ↩
- Ensign et al. (2018), a simulation using Oakland data; Mishler and Dalmasso (2022) on predictors fair when trained and unfair when deployed; Pagan et al. (2023) on feedback that can reduce bias. ↩
- Brynjolfsson, Chandar and Chen (2025), revised November 2025 and August 2026. ↩
- Humlum and Vestergaard (2025), NBER working paper, retitled in its March 2026 revision; the May 2025 version ruled out effects above 1%. ↩
- NASSCOM Strategic Review (February 2026), estimates, as recorded in the programme's facts log. ↩
- Ng et al. (2024): 70 developers, self-reported, without a control group. ↩
- Ding and Dafoe (2021) on electrification; Bick, Blandin and Deming (2024); Humlum and Vestergaard (2025); Acemoglu (2024), a model-based projection. ↩
Adapted for general readers from Working Paper P15 of the Embedded Systems Series (September 2026). The working paper carries the full argument, the provenance of every figure and the complete reference list. Model results are computed from stated models and are not estimates of any real system.