For a few years the standard explanation for AI chip shortages was lithography. ASML makes the only EUV scanners, so the story went, and everything downstream waits on them. By 2026 the binding constraint has moved. According to the book, the queue for an AI accelerator is set less by how many wafers a foundry can start than by a slot on TSMC's packaging line. The chip that comes out of the fab is no longer the finished article. It is one ingredient in a package that has become as hard to make, and as concentrated in its supply chain, as the silicon.
In this piece I want to explain why the industry moved performance gains from the transistor to the package, how the three main packaging architectures differ, what makes the physics of hybrid bonding so demanding, and why CoWoS capacity, not wafer starts, became the number people watch. Where the book flags a figure as directional or trade-press based, I have kept that caveat.
Why the package took over
For three decades, scaling meant shrinking the transistor. Below the 3 nm node that path gets steadily worse economically. Larger dies collect more defects per unit area. Lithography reticles cap the maximum die size at roughly 858 square millimetres. And the logic, memory and analog blocks inside a modern accelerator no longer benefit equally from the same leading-edge process. The industry's response was to stop building one giant chip and build several small ones, called chiplets, and to move the work of connecting them from the transistor layer to the package.
The defect argument is the core of it. On a large die, manufacturing defects scale with area, so doubling die size more than doubles the chance that at least one defect ruins the whole part. Split the same logic into smaller chiplets and each can be tested and sorted as a known-good die before it is committed to an expensive package. A defective memory stack or a marginal logic tile is discarded alone. Advanced packaging is as much a yield-economics strategy as a performance one: it keeps the yield benefits of small dies while recovering, through interconnect engineering, the bandwidth a monolithic die would have delivered for free.
Three answers to the same problem
The baseline is standard flip-chip ball-grid-array packaging. A die is soldered face-down onto an organic substrate with bump pitches of roughly 130 to 150 micrometres. That gives an interconnect density on the order of 100 bonds per square millimetre and an energy cost of roughly 1.5 to 2.5 picojoules per bit. It is fine for a desktop CPU or an automotive controller. It is nowhere near enough for a GPU that must move terabytes per second to a stack of High-Bandwidth Memory beside it.
2.5D silicon interposers. TSMC's CoWoS-S (Chip-on-Wafer-on-Substrate) is the dominant example. The logic die and HBM stacks sit side by side on a passive silicon interposer, a wafer with no active transistors, only through-silicon vias and several layers of fine copper wiring, with line and space widths reported down to about 0.4 micrometres. Because the interposer is patterned silicon, it inherits the reticle ceiling. CoWoS-S interposers are reported to top out at roughly 3.3 times the standard reticle field, on the order of 2,800 square millimetres. That caps the number of HBM stacks, typically six to eight, and the size of the logic die.
Embedded bridges. Intel's EMIB and TSMC's CoWoS-L are designed around that ceiling. Rather than one continuous interposer under the whole package, small silicon bridge chiplets, which TSMC calls local silicon interconnect, are embedded in an organic substrate, only under the die-to-die junctions that need fine routing. Silicon area shrinks, so the package is no longer bound to one reticle. CoWoS-L package sizes are reported to be approaching roughly 5.5 times reticle area by 2026, which is why the largest multi-die AI accelerators have moved toward bridges. The trade-off is a coarser bump pitch at the bridge, roughly 45 to 55 micrometres against 25 to 40 on a full CoWoS-S interposer, and somewhat lower density, at meaningfully lower silicon cost.
3D direct hybrid bonding. TSMC's SoIC and Intel's Foveros Direct remove the interposer altogether. Dies are stacked and bonded face to face through direct copper-to-copper diffusion, with no solder microbumps. Reported pitches run from roughly 3 micrometres down toward sub-1 micrometre on future roadmaps. Interconnect density exceeds 1,000,000 bonds per square millimetre, two to three orders of magnitude denser than a 2.5D interposer, and energy cost falls below 0.05 picojoules per bit.
| Platform | Bump pitch | Density | Energy | Representative use |
|---|---|---|---|---|
| Standard flip-chip BGA | 130 to 150 um | about 100 per mm2 | 1.5 to 2.5 pJ/bit | Desktop CPUs, automotive controllers |
| 2.5D interposer (CoWoS-S) | 25 to 40 um | about 1,500 to 2,500 per mm2 | 0.25 to 0.5 pJ/bit | AI GPUs such as Nvidia H100/B200, AMD MI300X |
| Embedded bridge (EMIB, CoWoS-L) | 45 to 55 um | about 800 to 1,200 per mm2 | 0.15 to 0.3 pJ/bit | Intel Xeon, data-centre CPUs, large multi-die AI packages |
| 3D hybrid bonding (SoIC, Foveros Direct) | under 3 um, roadmap sub-1 um | over 1,000,000 per mm2 | under 0.05 pJ/bit | AMD 3D V-Cache, next-generation AI logic stacking |
The book stresses that these are typical reported ranges across the industry, not a single vendor's guaranteed specification. They should be read as directional. Note also that the platforms are not ranked from worst to best. Each is a different point on a trade-off between package cost, achievable density and how much of the reticle and yield problem it solves. That is why all three are in production at once.
Underneath all three sits a standardisation layer, the Universal Chiplet Interconnect Express, or UCIe, an open die-to-die specification backed by Intel, TSMC, Samsung, AMD, Arm and Nvidia among others. It defines a common physical layer and protocol so chiplets from different vendors can in principle share a package. Its standard-package profile, for organic substrates at 100 to 130 micrometre pitch, targets about 28 gigabits per second per millimetre of die edge. Its advanced-package profile, for interposers and bridges down to 25 micrometre pitch, targets more than 1,300 gigabits per second per millimetre, with sub-0.25 picojoule per bit efficiency and latency under half a nanosecond.
Hybrid bonding: fusing copper at room temperature
Hybrid bonding gets its name because it joins two kinds of interface at once, dielectric to dielectric and metal to metal. Both wafers or dies are polished by CMP so that copper pads sit recessed a few nanometres below the surrounding dielectric, reported at about 2 to 5 nanometres, with surface roughness held under roughly 0.5 nanometres RMS. The surfaces are brought together at room temperature. The dielectric, usually silicon dioxide or SiCN, makes initial hydrophilic contact and forms a preliminary bond before any heat is applied.
A thermal anneal at roughly 300 to 400 degrees Celsius then completes it. The copper pads expand across the gap faster than the dielectric, close the recess and press the two copper surfaces into contact. Continued heating drives solid-state atomic interdiffusion across the interface, giving one continuous metallic bond with no solder, no intermetallic layer and no bump. Pitch is then limited by lithographic alignment accuracy, not by the mechanics of solder.
The bonding runs in an ISO Class 1 cleanroom because a single 100 nm particle trapped between two dies can void several bond pads at once. At these pitches, particle control is the process itself.
This precision is not free. Sub-nanometre flatness and single-digit-nanometre recess control across a 300 mm wafer are much harder than aligning a microbump. That is why hybrid bonding, though demonstrated in labs for years, has reached volume in a narrow set of products. The clearest proof is AMD's second-generation 3D V-Cache, which stacks extra L3 SRAM on top of CPU logic at a reported production pitch of about 9 micrometres. That is coarser than the sub-3 micrometre figures cited for the technology class. Today's highest-volume hybrid-bonding product runs at a pitch chosen for yield and cost, not at the physical ceiling. AMD's own reporting credits the approach with roughly a tenfold bandwidth improvement over conventional packaging. Roadmaps point to about 0.8 to 1.5 micrometres for next-generation products and later sub-1 micrometre, but the book is clear that this is a multi-year path, not a near-term default.
Known-good die logic applies with extra force. Hybrid-bonded dies fuse into one unit with no rework path, so every die must be tested and sorted before bonding. Test coverage before bonding is as much a gate on scaling as CMP and cleanroom limits are. Beyond V-Cache, the technique is commercially mature in 3D NAND, where it bonds peripheral logic wafers to memory-array wafers, and in CMOS image sensors. The frontier is logic-on-logic stacking, where Foveros Direct and SoIC aim to stack a whole compute die on a base die. That is a harder thermal problem. Removing heat from a die buried under another active die, with no microbump standoff to help, depends on backside power delivery and thermal interface materials that the book describes as still evolving. In a real sense the interconnect physics has outrun the industry's ability to cool what it enables.
The CoWoS crunch
The defining fact of advanced packaging in 2026 is that it, not front-end wafer fabrication, has become the binding constraint on AI accelerator supply. The book reports TSMC's CoWoS lines running at roughly 75,000 to 80,000 wafers a month in mid-2026, with targets of 115,000 to 140,000 by end of 2026 and 170,000 to 200,000-plus by 2027. Total 2026 CoWoS demand is estimated at close to one million wafers.
| Point in time | CoWoS capacity, wafers per month | Status |
|---|---|---|
| Mid-2026 | about 75,000 to 80,000 | Reported |
| End of 2026 | 115,000 to 140,000 | Target |
| 2027 | 170,000 to 200,000-plus | Target |
Even as capacity nearly doubles in about eighteen months, bookings have run ahead of it. Industry commentary describes CoWoS-S and CoWoS-L as effectively sold out, with lead times of roughly 52 to 78 weeks, over a year from order to packaged part. I want to be careful here, as the book is. Those lead-time figures vary by source, were flagged in the trade-press reporting as directional, and no TSMC-disclosed number was found. They illustrate the scale of the backlog. They are not a commitment.
A large share of that queue belongs to one customer. Nvidia is reported, across several independent trade-press estimates and not a single confirmed figure, to have booked roughly half to as much as 60 percent of TSMC's CoWoS capacity through 2026 and 2027. That effectively sets the allocation queue for every other accelerator and networking-chip designer. It is a market-structure risk separate from the physical capacity limit: a firm with a funded, taped-out design can still wait on a packaging slot that another roadmap has already claimed. The same applies to hyperscaler custom-silicon programmes and to networking vendors that need fine-pitch packaging for switch and optical-engine silicon. For them the CoWoS queue, not engineering execution, is a throttle on how fast they can grow.
Three chokepoints in series
The more useful comparison with lithography is structural. Lithography has one chokepoint. Advanced packaging has three in series, spread over three countries, and none has a qualified second source at scale.
Ajinomoto's roughly 95 percent share of ABF, the build-up film dielectric used in advanced package substrates, sits in Japan. TSMC's dominant share of CoWoS-class capacity sits in Taiwan, with one Morgan Stanley estimate cited in trade press putting it above 80 percent of global CoWoS-class wafer demand. HBM is a three-firm oligopoly of SK Hynix, Samsung and Micron concentrated in South Korea, with SK Hynix estimated at roughly 50 to 62 percent depending on quarter and source. A package cannot exist without clearing all three, and the packaging slot is of limited use without an HBM allocation already secured. That is a tighter concentration profile than lithography ever had.
The substrate layer compounds the squeeze. ABF substrate lead times were reported at 16 to 24 weeks in mid-2026, with the supply-demand gap estimated at about 10 percent in the second half of 2026, roughly 21 percent in 2027 and possibly beyond 40 percent by 2028. Ajinomoto is reported to be raising ABF film prices by around 30 percent, with an estimated 3 to 6 percent pass-through to overall AI substrate cost.
Geography adds risk. Because CoWoS-class capacity sits overwhelmingly in Taiwan, and Taiwanese policy is reported to require that cutting-edge process and packaging work remain onshore, a Taiwan Strait disruption would leave no near-term alternate source. The book quotes trade-press estimates, explicitly caveated as illustrative commentary and not a validated forecast, of roughly a 20 percent probability for a serious escalation scenario, with lead-time extensions beyond six months and 25 to 35 percent price rises on critical nodes if it happened. I treat that number as a sizing of the exposure, not a prediction.
Foundry in-house, or OSAT?
Advanced packaging used to sit downstream of the foundry, in the hands of outsourced assembly and test firms. For the highest-value 2.5D and 3D work that boundary has dissolved. CoWoS-class interposers and hybrid bonding demand the same cleanroom class, process control and yield learning as front-end lithography, so leading foundries pulled the highest-margin work in-house.
| Player | Advanced-packaging role | Scale and recent moves |
|---|---|---|
| TSMC | In-house 3DFabric: CoWoS-S, CoWoS-L, SoIC | Over 80% of CoWoS-class demand (estimate); expanding Tainan, Chiayi, Zhunan |
| Samsung | I-Cube, X-Cube, R-Cube, H-Cube | Fast follower; expanding packaging and test in Vietnam, reported about US$4B |
| Intel | Foveros (3D), EMIB (bridge); EMIB-T, EMIB 3.5D in development | US build-out in Arizona and New Mexico |
| ASE | Largest OSAT; overflow and complementary capacity | 44.6% of OSAT revenue; six new facilities from 2026; quotes reportedly up 20%+ |
| Amkor | Second-largest OSAT; licensed extension of TSMC CoWoS | June 2026 ten-year TSMC agreement for local CoWoS capacity |
The book splits the field into two tiers. The top tier, with sub-40 micrometre pitches, TSV interposers and hybrid bonding, is treated by TSMC, Samsung and Intel as an extension of front-end manufacturing, with the same capital intensity and increasingly the same customer relationships. A customer who trusts TSMC with a 3 nm wafer has an obvious reason to trust it with the packaging. The tier below, conventional flip-chip, wirebond and fan-out plus final test, remains the OSAT core, where ASE, Amkor, SPIL, JCET and Powertech compete on cost and capacity.
Amkor's June 2026 agreement is the most telling move. TSMC will buy local CoWoS packaging and test capacity from Amkor's facilities, which makes Amkor a licensed extension of TSMC's ecosystem. TSMC widens its footprint without ceding process technology or the customer relationship. That is a different thing from an OSAT winning independent CoWoS business, and it may be the template for others.
One long-horizon release valve sits outside the foundry-versus-OSAT axis: glass-core substrates, developed most visibly by Corning. Glass offers a more consistent thermal expansion coefficient than ABF, reducing warpage, and supports laser bonding and debonding, lowering chip-breakage risk. Intel cited its Clearwater Forest product as a January 2026 production milestone on glass. The book places readiness at roughly 4 to 5 on the technology scale, with standards still being defined and pricing at two to five times organic substrate, with parity not expected until about 2028. It is a real alternative to the ABF chokepoint, but not within the 2026 to 2027 window of this crunch.
Where India sits
The book is direct on this. India has no confirmed presence at the interposer, TSV, micro-bump or hybrid-bonding tier. Its four live or near-live packaging investments, Micron at Sanand, Kaynes at Sanand, CG Semi at Sanand and Tata Electronics at Jagiroad in Assam, sit at the conventional assembly, test and packaging tier, well downstream of the CoWoS-equivalent chokepoint. No government or private roadmap targeting silicon-interposer capability was identified, and the book marks that as an evidence gap and not a confirmed absence. I cover that back-end tier in the companion piece on OSAT economics.
What I take away
The move to chiplets is a yield decision first. Small, tested dies beat one large die, and the price of that choice is that the interconnect between them has to do the work the die used to do for free.
The three architectures are complements. Interposers give the highest density that fits under a reticle-limited silicon area. Bridges give larger packages at lower silicon cost. Hybrid bonding gives more than a million bonds per square millimetre but demands sub-nanometre flatness, an ISO Class 1 cleanroom and known-good dies, and its thermal problem is unsolved.
The bottleneck shifted because packaging capacity is finite, pre-allocated years ahead and concentrated. Three chokepoints in three countries, one customer holding roughly half or more of the leading line, and lead times reported over a year: that is a tighter structure than the lithography story, even if the exact lead-time and share figures are directional.
For anyone planning supply, the practical point is that a funded design is not a shippable design until it has a substrate, an interposer slot and an HBM allocation. And for countries outside this chain, the honest position is that packaging at this tier is not yet within reach, and the first rung is the conventional back end.