AnyLearn
All lessons
Historyintermediate

Concentration and resilience: what it takes to de-risk a chip supply chain

Mapping the single points of failure in semiconductor production, why building fabs elsewhere is necessary but nowhere near sufficient, the real toolkit of resilience strategies with their costs, and a transferable method for analyzing any concentrated supply chain.

Updated · AI-authored, review-gated · how lessons are made

Not signed in: your progress and quiz score won't be saved.
Progress1 / 7

The risk register, stated plainly

The previous lessons established the structure; this one asks what happens when parts of it fail, and what failure-proofing actually costs.

The concentration facts worth keeping on one card:

  • Advanced logic: on the order of 90% produced by one firm, with the most advanced fabs on one island (Taiwan) that sits in an earthquake zone and a geopolitical flashpoint.
  • Advanced packaging (the stacking and connecting of dies that AI accelerators depend on): dominated by the same region.
  • EUV lithography: one vendor (Netherlands), itself dependent on single-source optics (Germany) and lasers (Germany).
  • Critical materials: EUV photoresists, mask blanks, and the majority of wafers from a handful of Japanese firms; specialty gases with concentrated sources (neon refining, for example, was historically concentrated in Ukraine, a fragility exposed in 2022).

The risk types differ: natural disaster (earthquake, fire, drought, a single fab uses millions of liters of ultrapure water daily), accidents (individual plant fires have historically moved global prices of specific chemicals and memory), pandemics and logistics failures, and armed conflict, the scenario around Taiwan that motivates most current policy. A risk register like this is standard practice for any industry; what makes semiconductors unusual is how few nodes carry how much consequence.

Full lesson text

All 7 steps on one page, for reading, reference, and search.

Show

1. The risk register, stated plainly

The previous lessons established the structure; this one asks what happens when parts of it fail, and what failure-proofing actually costs.

The concentration facts worth keeping on one card:

  • Advanced logic: on the order of 90% produced by one firm, with the most advanced fabs on one island (Taiwan) that sits in an earthquake zone and a geopolitical flashpoint.
  • Advanced packaging (the stacking and connecting of dies that AI accelerators depend on): dominated by the same region.
  • EUV lithography: one vendor (Netherlands), itself dependent on single-source optics (Germany) and lasers (Germany).
  • Critical materials: EUV photoresists, mask blanks, and the majority of wafers from a handful of Japanese firms; specialty gases with concentrated sources (neon refining, for example, was historically concentrated in Ukraine, a fragility exposed in 2022).

The risk types differ: natural disaster (earthquake, fire, drought, a single fab uses millions of liters of ultrapure water daily), accidents (individual plant fires have historically moved global prices of specific chemicals and memory), pandemics and logistics failures, and armed conflict, the scenario around Taiwan that motivates most current policy. A risk register like this is standard practice for any industry; what makes semiconductors unusual is how few nodes carry how much consequence.

2. Why efficiency built exactly this

The fragility is not an accident of neglect; it is the purchase price of efficiency, and seeing that mechanism clearly is the lesson's core.

Recall the economics: fabs reward utilization and scale (lesson 1); process layers reward specialization and experience (lessons 1-2). For forty years, every actor optimizing locally, cheaper chips, better yields, faster iteration, pushed each layer toward its most efficient configuration, which is a small number of giant, specialized, geographically clustered facilities:

  • Clustering fabs in one region shares talent pools, supplier networks, universities, and logistics (the mechanism behind every industrial cluster since textile towns).
  • Single-sourcing a material from its best producer maximizes purity and minimizes cost.
  • Letting one firm win lithography concentrated the R&D spend the technology needed to exist at all.

Each choice was individually rational; their sum is a system with minimal redundancy. This is a general law of optimized systems: efficiency removes slack, and slack is what absorbs shocks. The corollary matters for policy analysis: resilience is not free repair work around the edges; it means deliberately buying back slack that the market rationally eliminated, and someone (taxpayers, customers, shareholders) pays the carrying cost, permanently.

3. "Just build fabs elsewhere": the necessary-but-insufficient move

The instinctive fix, subsidize fabs in new locations, is where every major economy started: multi-tens-of-billions programs in the United States (the CHIPS Act's $52B+), the European Union, Japan, India, and others, with headline fabs under construction or ramping in Arizona, Ohio, Germany, and Kyushu.

What the money buys, and what it doesn't:

  • It buys buildings and tools. Capital was never the binding constraint for incumbents; it is the entry ticket, roughly $25B per leading-edge fab (lesson 1's arithmetic), before a single wafer ships.
  • It does not buy the ecosystem. A fab imports its chemicals, gases, spare parts, and consumables continuously; if those still come from the original clusters, the location of the building changes while the dependency graph barely moves. Studies of new-site economics consistently find meaningful cost premiums over established clusters (commonly cited in the tens of percent), driven by construction, talent scarcity, and thin local supplier bases.
  • It does not buy yield culture quickly. Ramp timelines at new sites have repeatedly slipped versus home-cluster fabs; the tacit-knowledge transfer (engineers rotating from the home fab, sister-fab copying of process recipes) is the actual bottleneck, and it moves at the speed of people.
  • It changes the risk profile, not the risk level, at first. A second site for older nodes diversifies real production; a satellite fab whose leading-edge recipes, engineers, and materials still flow from the original cluster mostly adds a dependent branch, until, over years, the branch roots.

None of this makes diversification pointless; it calibrates it: fabs-elsewhere is a decade-scale project whose early years deliver far less independence than the ribbon-cuttings suggest.

4. The full resilience toolkit, priced

Fabs-elsewhere is one row of a larger menu. The honest version of the menu shows what each strategy protects against, and what it costs:

StrategyProtects againstStanding cost / limit
Geographic fab diversificationRegional disaster, conflict$10s of billions per site; ecosystem lag; ~10-30% cost premium
Second-sourcing materials/equipmentSingle-supplier failureQualification takes years; purity/quality risk; splits the volume that funds the best supplier
Strategic stockpilesShort interruptions (months)Carrying cost; chips and chemicals age; wrong-mix risk
Design flexibility (multi-fab qualified designs)Losing one fabExtra engineering per design; performance compromises
Mature-node onshoringAuto/industrial chip shocksCompetes against depreciated overseas fabs; needs demand guarantees
Packaging diversificationThe most concentrated back-end stepLabor economics resist relocation; automation helps slowly
Demand-side buffering (inventory, contracts)Price and lead-time spikesWorking capital; bullwhip amplification if everyone does it at once

Two readings. First, strategies are layered, not alternative: stockpiles cover the first months, second sources the first years, new clusters the decade. Second, every row is a standing premium, resilience is an insurance product, and the recurring policy failure is buying it once (a subsidy cycle, a stockpile) and letting it lapse when prices normalize, the well-documented boom-bust pattern of attention to this industry.

5. Case study: the 2020-2023 chip shortage, mechanism by mechanism

The pandemic-era shortage is the cleanest natural experiment on record, and notably, it barely involved the leading edge. The chips that halted car production lines were mature-node parts costing a few dollars each. The mechanism chain:

  1. A demand forecast error. Automakers, expecting a long slump, cancelled chip orders early in 2020. Consumer-electronics demand instead surged (remote work), and foundries reallocated the freed capacity instantly, utilization pressure guarantees that (lesson 1).
  2. A rigid re-entry queue. When car demand rebounded within months, the capacity was gone, and mature-node capacity has no slack by design: those depreciated fabs are precisely the ones nobody builds new. Lead times stretched past a year.
  3. Qualification rigidity. Car makers could not simply buy substitute chips: automotive parts require years-long safety qualification, so even functionally similar chips from other lines were unusable in the short run. The dependency was on specific qualified parts, far narrower than "chips."
  4. The bullwhip. Every buyer, burned once, began over-ordering and double-booking, inflating apparent demand, which extended shortages and then produced the equally sharp glut of 2023 when real demand normalized.

Estimated cost to the auto industry alone: on the order of $200 billion in lost 2021 revenue, from parts worth a rounding error of that. The transferable reading: fragility concentrated not where technology was most advanced, but where margins were thinnest, capacity oldest, and substitution hardest. Resilience analysis that only watches the leading edge misses where the last crisis actually happened.

6. The interests are not aligned, and that shapes outcomes

Resilience analysis goes wrong when it treats "the supply chain" as one actor with one goal. The load-bearing actors want different things:

  • The concentrated producers benefit from concentration; it is their moat. A leading foundry diversifying its own geography does so for customer reassurance and subsidy capture, while keeping the leading edge, and the pricing power, at home. Observed behavior matches: newest nodes ramp at the home cluster first, satellites follow years later.
  • Host governments of the clusters hold a strategic asset precisely because the world depends on them (the "silicon shield" logic); complete de-risking by others would spend that asset. Cooperation is therefore real but bounded.
  • Chip customers want resilience they do not have to pay for; given a price premium for a diversified source versus a cheaper concentrated one, procurement has historically chosen cheap, which is how the concentration formed. Post-shortage behavior (multi-sourcing clauses, inventory buffers) fades measurably as memories fade.
  • Subsidizing governments need visible, attributable wins on electoral timescales, which biases spending toward ribbon-cuttable fabs over unglamorous, higher-leverage targets like materials qualification, packaging, and workforce pipelines.

The prediction this misalignment supports is structural, not speculative: diversification proceeds, but lags the leading edge persistently, because every actor's incentives point that way. Reading any resilience announcement means asking: which actor is paying, which is de-risked, and which dependency actually moved?

7. A method you can reuse on any supply chain

The semiconductor case yields a five-step analysis that transfers to batteries, pharmaceuticals, rare earths, undersea cables, or cloud infrastructure:

  1. Draw the real dependency graph. Not the shipping map, the capability map: for each layer, who can actually perform it at the required grade, and in which jurisdictions? (Lesson 1's table; the chokepoint is rarely the famous firm.)
  2. Classify each chokepoint's moat. Capital-scale (buildable with money and years), ecosystem/tacit-knowledge (decades), or physics/IP (may be effectively unreplicable on relevant timescales). The moat type sets the realistic de-risking horizon.
  3. Price the failure modes. Duration and breadth: a fire is months and one material; a regional conflict is years and everything colocated. Match remedies by horizon: stockpiles ↔ months, second sources ↔ years, new clusters ↔ decades.
  4. Check who pays the standing premium. Resilience that no actor has a durable incentive to fund will decay; look for structures (regulation, long-term contracts, mandated reserves) that lock the premium in.
  5. Expect the system to adapt against you. Controls beget design-out and indigenization (lesson 3); subsidies beget subsidy-shopping; buffers beget bullwhips. Analyze the second move, not just the first.

Applied to chips, the method explains both headline facts of the current era: why so much money is moving (the failure modes are priced in trillions), and why the map changes so slowly (the deepest moats are made of time, and the incentives to keep them are strong). That is the equilibrium this path set out to explain, from wafer economics to policy, one mechanism at a time.

Check your understanding

The lesson ends with a 5-question quiz. Take it in the player above to see your score.

  1. Why is the semiconductor supply chain's fragility described as "the purchase price of efficiency"?
    • Fabs deliberately avoid backup systems to cut insurance costs
    • Regulators required concentration for safety reasons
    • Efficiency is unrelated to fragility; the chain is fragile by accident
    • Decades of locally rational optimization (scale, specialization, clustering, single-sourcing) removed the slack that would absorb shocks
  2. Why does subsidizing a new fab in a new region deliver less independence than it appears to, especially early on?
    • New fabs are legally barred from leading-edge production
    • The building and tools arrive, but materials, spare parts, recipes, and experienced engineers still flow from the original clusters for years
    • Subsidies can only be spent on mature nodes
    • New regions lack electricity for fabs
  3. How should the resilience strategies (stockpiles, second sources, new clusters) be combined?
    • Choose exactly one to avoid redundancy
    • New clusters first, since they are fastest
    • Layered by time horizon: stockpiles cover months, second-sourcing covers years, new clusters cover the decade
    • Stockpiles alone suffice if large enough
  4. Which incentive misalignment helps explain why leading-edge production diversifies so slowly?
    • Concentrated producers' moat IS the concentration, so they ramp newest nodes at home and satellite fabs later; customers historically choose cheaper concentrated supply
    • No government has offered subsidies for new fabs
    • Chip customers uniformly pay premiums for diversified supply
    • Host governments of existing clusters push their key firms to move abroad quickly
  5. In the transferable five-step method, why must you "analyze the second move, not just the first"?
    • Because supply chains change too slowly to matter
    • Optimized systems adapt against interventions: controls beget design-out, subsidies beget subsidy-shopping, buffers beget bullwhip effects
    • Because the first move is always secret
    • Regulators require two rounds of public comment

Related lessons