---
title: "The Data Center Bubble: A Layered Reading of Risk in the Artificial Intelligence Economy"
slug: data-center-bubble-ai-2026
language: en
description: "France Épargne research paper: the AI boom is real, the bubble sits in data centers. Layered risk map, stranded assets, GPU to ASIC shift, 2026 to 2030 scenarios."
keywords: [AI bubble, data center bubble, hyperscaler capex 2026, stranded assets, GPU to ASIC transition, edge inference, scaling laws, AI infrastructure risk]
tags: [Intelligence Artificielle, Centres de Données, Semi-conducteurs, Bulle Spéculative, Infrastructure, Hyperscalers, Nvidia, GPU, ASIC, Marchés Financiers]
canonical: "https://www.france-epargne.fr/research/en/data-center-bubble-ai-2026"
author: France Épargne Research
classification: Strategic Intelligence. Public Distribution
publishedAt: "2026-08-03T00:00:00.000Z"
updatedAt: "2026-08-03T21:04:29.316Z"
readingTimeMinutes: 205
---
> Intelligence Boom, Data Center Overbuild, Hardware Paradigm Shift

# The Data Center Bubble: A Layered Reading of Risk in the Artificial Intelligence Economy

## Abstract

The question "is AI a bubble?" cannot be answered as posed, because "AI" names at least eight distinct economic layers: foundation models, semiconductors, cloud platforms, data centers, energy, applications, wrappers, and distribution. Each layer carries its own cost structure, its own revenue base, and its own failure mode, and the June 2026 econometric literature finds bubble signals concentrated in some layers while others price against verifiable fundamentals.[^1] This paper defends a decomposed thesis: artificial intelligence usage is durable and compounding, the application excess purges itself continuously with each model release, and the genuine systemic risk sits in centralized compute infrastructure, whose economics assume the permanence of a specific hardware paradigm. History supports the decomposition. The British railway mania of the 1840s halved shareholder wealth while financing a national network that operated for a century; the telecom crash of 2001 to 2002 destroyed roughly two trillion dollars of market value while internet traffic kept doubling every year through the collapse. The pattern, formalized by Carlota Perez as the sequence of installation, crash, and deployment, predicts that a crash in AI infrastructure would be a phase transition inside a technology boom rather than the end of one. The mid 2026 data already exhibit the signature decoupling: Google reported 3.2 quadrillion tokens processed per month in May 2026, a roughly 330 fold increase in two years, while in July 2026 the Philadelphia Semiconductor Index fell 28.6 percent from its June peak and the sector's flagship leveraged fund liquidated. We map the risk gradient across the stack, anchored in post May 2026 figures: 725 billion dollars of hyperscaler capital expenditure against an AI revenue base near one tenth of annual industry spend, 30 to 50 percent of planned American data centers delayed or canceled, and adoption metrics accelerating through the financial stress. The conclusion for allocators is operational. Underwrite each layer separately, treat aggregate "AI exposure" as a category error, and expect the coming correction, if it arrives, to be misread as the death of a technology whose usage curve does not flinch.

## 1. Introduction: deconstructing "the AI bubble"

### 1.1 A July that proved both camps right

The last week of July 2026 delivered, within four trading days, the strongest evidence for the bubble thesis and the strongest evidence against it. On July 28, Alphabet's shares fell 7 percent after its capital expenditure report, dragging Amazon, Meta, and Microsoft down with it, as investors scrutinized shrinking cash piles set against uncertain returns.[^2] By July 31, the Philadelphia Semiconductor Index had fallen 28.6 percent from its June 22 peak and the Morgan Stanley Momentum TMT Index had dropped 53.5 percent in a single month.[^3] Situational Awareness LP, the sector's flagship leveraged vehicle, collapsed from 45 billion dollars to roughly 10 billion dollars in late July and sold assets to Citadel in a forced liquidation.[^4] For a reader of headlines, the AI bubble was bursting.

The same summer produced the opposite dataset. At Google I/O on May 20, 2026, the company reported processing 3.2 quadrillion tokens per month across its products and interfaces, up from 480 trillion in May 2025 and 9.7 trillion in May 2024: a 330 fold expansion of usage in twenty four months, still compounding at roughly 7 times per year.[^5] OpenRouter, an independent routing layer that aggregates demand across model providers, crossed 25 trillion tokens per week in the same month, 5 times its volume of six months earlier.[^105b] Whatever was collapsing in July 2026, it was not the demand for machine intelligence.

Both observations are correct, and the tension between them dissolves once the object under analysis is specified. The July drawdown hit semiconductor equities, momentum portfolios, and a leveraged fund whose thesis required centralized compute buildout to continue on schedule. The token curve measures usage of the technology itself. These are different assets, in different layers, with different risk profiles, and the central argument of this paper is that no serious capital allocation decision can be made until the question "is AI a bubble?" is replaced by the question "which layer of the AI economy is mispriced, against which revenue base, on which timeline?"

### 1.2 What people mean when they say "AI bubble"

The aggregate framing dominates public discourse. In November 2025, National Public Radio asked whether an AI bubble was brewing by applying the Shiller price to earnings ratio, a single valuation gauge, to "AI stocks" as one object.[^6] In July 2026, coverage of a Bank for International Settlements systemic risk study announced that the "AI boom outgrows every tech bubble in history," again treating the phenomenon as one measurable mass.[^7] The framing is understandable: the numbers involved are aggregate and enormous. AI industry spending runs at roughly 400 billion dollars per year against 50 to 60 billion dollars of annual AI revenue.[^8] The four largest hyperscalers plan approximately 725 billion dollars of capital expenditure in 2026, an increase of 77 percent in a single year.[^9] That capital expenditure (**capex**, spending on long lived physical assets) is projected to consume about 94 percent of hyperscaler operating cash flow across 2026 and 2027, and the combined free cash flow of Amazon, Google, Meta, and Microsoft is forecast to shrink 43 percent between the fourth quarter of 2024 and the first quarter of 2026.[^10]

A spending to revenue gap of that magnitude is a legitimate object of alarm. The error is in the address label. The gap does not sit uniformly across "AI." It sits between specific balance sheets (the entities financing physical infrastructure) and a specific revenue base (the entities selling intelligence), and those are largely different companies, or at minimum different divisions with different economics inside the same companies.

### 1.3 Why the aggregate label fails analytically

Three bodies of work converge on the same conclusion. First, the platform economics literature has held since 2010 that digital ecosystems are layered modular architectures: devices, networks, services, and contents form loosely coupled strata, each with independent innovation dynamics and independent competitive logic.[^11] Second, current industrial organization research models the AI value chain the same way. A 2025 study in Industrial and Corporate Change analyzes the chain as closely interconnected but analytically distinct layers (infrastructure, foundation models, generative applications, users), each with its own competitive dynamics and its own policy levers.[^12] The Computer and Communications Industry Association states the practical version plainly: "AI is not a single monolithic market; it is a family of technologies operating at different but related layers, from chips and data infrastructure to algorithms and end-user applications."[^13]

Third, and most directly, the first multi method econometric evaluation of the bubble question, published as a working paper in June 2026 by Wang and Chen, reaches a differentiated verdict: capex has accelerated faster than monetization "in some layers," while adoption and productivity fundamentals remain genuine, leading the authors to describe AI as "a real technological revolution with localized bubble dynamics" rather than either a pure mania or a bubble free productivity miracle.[^1] Practitioners have begun operationalizing the same insight: the ClearTake AI Bubble Index is computed separately for each value chain layer, showing where speculative pressure concentrates rather than issuing one verdict for the whole economy.[^14]

An analyst who prices Nvidia, a data center landlord, and a thin chat application as one "AI exposure" bucket is therefore averaging assets whose correlation is contingent on financing arrangements rather than on shared fundamentals. The thematic funds and index products built on that single bucket are among the instruments most exposed to the correction this paper anticipates, precisely because they will transmit a shock in one layer to holdings in every other.

### 1.4 The thesis, stated in full

The paper defends nine connected claims, presented here openly so the reader can hold the argument against the evidence as it accumulates.

1. The phrase "AI bubble" is analytically empty until the speaker specifies which layer is supposedly mispriced. The ecosystem decomposes into foundation models, chips, cloud, data centers, energy, applications, wrappers, developer tooling, and distribution, and these layers carry fundamentally different risk profiles.[^1]

2. Historical technology bubbles destroyed weak financial vehicles while the underlying technology kept compounding. Railway shareholders lost roughly half their capital between 1846 and 1850; Britain kept the railways. Telecom investors lost on the order of two trillion dollars by 2002; the internet kept doubling. The tulip analogy misleads whenever it is used to imply that the underlying object loses its utility.[^15]

3. AI usage is durable and expanding, and the demand for intelligence behaves as a moving frontier: each solved capability raises ambition toward the next task, so current use cases systematically understate the future market. The token consumption record through May 2026 is consistent with this dynamic.[^5]

4. Today's dominant applications (text, chat, code assistance, thin software as a service) are the oil lamp phase of a general purpose technology. The equivalents of aviation fuel, plastics, and pharmaceuticals (AI for science, robotics, medicine, materials) have not been fully discovered, which biases every present tense market sizing downward.

5. The wrapper layer is genuinely bubble like, and it purges continuously: each frontier model release absorbs the weakest applications, so the excess deflates release by release instead of accumulating into a single crash. Wrappers meanwhile function as market research for the laboratories, and many AI founders are economically customers of the platforms rather than owners of durable businesses (evidence grade: B).[^16]

6. Large platform valuations are not automatically irrational. They must be evaluated against the hypothesis that AI becomes a utility scale spending category, benchmarked against telecom average revenue per user for consumers and software spend per employee for enterprises.

7. The genuine systemic risk sits in centralized compute infrastructure, whose buildout prices in the permanence of the current paradigm: brute force scaling, training bound to graphics processing units, centralized inference. Algorithmic efficiency, model saturation, and hardware specialization each threaten that assumption.

8. Training demand and inference demand must be separated. Token consumption can grow by orders of magnitude while the economics of late, long horizon, GPU centric data center projects collapse, because demand for intelligence and demand for centralized compute capacity are different variables (evidence grade: B).[^5]

9. If the infrastructure layer cracks, the event will be publicly misread as "the AI bubble bursting," when it will in fact mark a hardware paradigm transition inside a genuine intelligence boom. After the crash, deployment accelerates on cheaper infrastructure, following the historical template (evidence grade: B).[^17]

Claims 1 and 2 rest on directly relevant published evidence and are graded A in our evidence base. Claims 5, 8, and 9 combine documented mechanisms with forward looking projection and are asserted as graded hypotheses, with the grade shown where the stakes justify it.

### 1.5 The strongest objection, taken first

The best argument for the aggregate label deserves its hearing before the decomposition proceeds. The Bank for International Settlements systemic risk work, as reported in July 2026, argues that circular financing arrangements recouple the layers financially even where their operating risks differ: hyperscalers take equity stakes in laboratories that spend the proceeds buying compute back from the same hyperscalers, creating a closed valuation loop in which a shock at one node propagates to all.[^7] Bloomberg's reconstruction of the deal graph among Microsoft, OpenAI, and Nvidia documents how much recorded AI revenue circulates inside the ecosystem rather than arriving from end customers.[^10]

The objection is partially correct and fully compatible with the thesis. Circular financing raises the correlation of financial outcomes across layers; it does not equalize their operating fundamentals. When the telecom crash arrived in 2001, vendor financing had similarly coupled equipment makers to carriers, and the coupled equities fell together. What the coupling could not do was make internet usage fall: traffic kept doubling roughly annually straight through the collapse of the balance sheets built on false traffic claims.[^15] Financial contagion determines who goes bankrupt in a correction. Operating fundamentals determine what exists five years later. The layered analysis is precisely the tool that separates those two questions, and the circular financing evidence strengthens the case for performing it, because it identifies the transmission channels through which a data center shock would masquerade as a technology shock.

### 1.6 Why the misreading is predictable

Narrative economics gives the misreading a mechanism rather than leaving it as a complaint about journalism. Robert Shiller's account of how markets process crises holds that the public understands economic events through viral simplified stories, with the news media acting as the key generator and disseminator of those narratives; contagion of the story, rather than the underlying mechanics, determines how an event is collectively understood.[^17] Empirical work in behavioral finance tracking media coverage across the dot-com mania, the global financial crisis, and the COVID crash confirms that extreme market periods crystallize into a small set of emotionally loaded narratives.[^18]

The 2001 precedent shows the compression at work. The telecom collapse was narrated as the internet having been overhyped, while Andrew Odlyzko's traffic data showed usage doubling every year through the crash; the failure lived in carrier balance sheets and in an industry legend that traffic doubled every 100 days, and the narrative erased the distinction.[^15] Sophisticated observers already resist the compression today. The Man Group's February 2026 analysis locates the bubble explicitly in infrastructure financing and states the layer distinction in one sentence: "AI technology is transformative and here to stay. But the inflated financial architecture supporting it may be unsustainable."[^16] The mass narrative channel, however, runs on the aggregate label, and the July 2026 market action, in which a capex disclosure by one company moved the entire "AI trade" as a block, shows the aggregate reflex operating in real time.[^2] We therefore assert, as a graded hypothesis (evidence grade: B), that a data center correction will be publicly processed as "the AI bubble bursting," and that the mispricing this misreading generates in the technology's genuinely compounding layers will be among the larger alpha opportunities of the cycle.[^17]

### 1.7 Route map

Section 2 establishes the historical base rates: what tulip mania actually was, what railway and telecom investors lost, what the technologies delivered afterward, and why the Perez installation and deployment framework organizes all of it. Section 3 decomposes the AI economy into its layers and anchors each layer's size and risk profile in post May 2026 figures. Later sections develop the wrapper purge and the founder as customer economy (Sections 4 and 5), the expanding demand frontier (Section 6), the general purpose technology argument and the valuation benchmarks (Sections 7 and 8), the scaling paradigm and its limits (Section 9), the infrastructure trap and stranded asset timing (Sections 10 and 11), the steelmanned counterarguments (Section 12), and the synthesis with named winners, threatened incumbents, and sleepers (Section 13).

[^105b]: Business Wire, OpenRouter announces 113 million dollar Series B led by CapitalG, May 26, 2026, https://www.businesswire.com/news/home/20260526953416/en/

## 2. What history actually teaches

### 2.1 The spine: installation, crash, deployment

Carlota Perez's study of five technological revolutions supplies the framework this section hangs on. In her account, every major technology surge divides into an **installation phase**, in which speculative financial capital funds the frenzied overbuilding of the new infrastructure, and a **deployment phase**, in which productive capital exploits the installed base for decades of broad growth. Between the two sits a crash: the turning point at which the financial vehicles created during the frenzy are destroyed and asset prices reconnect with earning capacity. Perez's own formulation assigns the crash a productive function: "Financial capital then acts as the agent of massive creative destruction," after which the recoupling of financial and production capital enables sustained growth.[^19] Commentary applying the framework to the 2026 AI buildout has become common; a July 2026 reading summarizes the template as "Every technological revolution runs the same financial script: installation, frenzy, crash, then the golden age."[^20]

The framework earns its place through a specific empirical regularity, documented case by case below: in each historical episode, the crash destroyed a class of financial claims while leaving the technology's diffusion curve essentially untouched. Investors and the technology had different outcomes because they were exposed to different things. The investor owned a claim on future cash flows priced at the peak of the frenzy; the adopter bought service at marginal cost from overbuilt capacity. Overcapacity plus competition transferred the technology's surplus from the capital that installed it to the users who deployed it.[^21] Holding that mechanism in view, we can take the five standard episodes in order, extracting from each what investors lost, what the technology delivered afterward, and what the pattern implies for AI.

### 2.2 Tulips: the archetype that fails on inspection

Tulip mania is the reflexive analogy for any suspected bubble, and it is the least informative of the set. Peter Garber's revisionist study, Famous First Bubbles, reexamined the Dutch archives and concluded that most of the episode is explicable by market fundamentals rather than collective madness.[^22] Three findings matter. First, the famous prices attached to rare, virus induced bulb varieties whose supply was genuinely fixed in the short run, and steep price declines for prized new varieties were the normal life cycle of the market: Garber documents that the most expensive hyacinth varieties of the early nineteenth century likewise fell to 1 to 2 percent of peak value within thirty years, with no accompanying legend of madness.[^22] Second, the notorious final months of the mania ran on forward contracts in taverns, agreements in which little or no money changed hands before the crash voided them; Garber characterizes the late common bulb trade as closer to a winter drinking game among a plague stricken population than to a leveraged financial mania.[^22] Third, the measurable economic fallout was small: no wave of bankruptcies, no credit contraction, no depression traceable to bulbs. Garber adds a contemporary calibration for readers who find seventeenth century bulb prices self evidently mad: prototype bulbs of newly bred lily varieties have changed hands in the modern Dutch market for around one million guilders, because a novel variety's entire future supply descends from a handful of initial specimens.[^22]

The relevance to AI is diagnostic. When a commentator reaches for tulips, they import two implications: that the priced object had no underlying utility, and that the crash left nothing behind. Neither implication survives Garber's evidence even for tulips themselves (the Netherlands built a permanent global flower industry), and both fail completely for infrastructure technologies. Tulip mania, whatever it was, financed no canals, rails, cables, or generating stations. It is therefore the wrong template for evaluating a buildout, and its persistence in the AI debate measures the gap between narrative convenience and analytical use. The episodes that do resemble the present, examined next, share a feature tulips lack: the speculative excess purchased physical infrastructure that outlived its financiers.

### 2.3 Railway mania: ruined shareholders, a working network

The British railway mania of the 1840s is the cleanest instance of the decoupling. Andrew Odlyzko's study of the episode, built on the financial press and share registers of the period, documents the scale on both sides. On the investment side, roughly 1,000 new railway companies were floated, and railway share prices fell approximately 50 percent between 1846 and 1850. The losses were amplified by the era's financing structure: shares were partly paid, with investors liable for calls on the roughly 90 percent of face value not yet paid in, so the decline arrived together with demands for fresh capital into a falling market.[^23] On the technology side, the mania financed the basic structure of the British railway system, a network that operated, essentially as laid out, for the following century.[^23]

Gareth Campbell's work on the mania adds a discipline against easy moralizing. Examining the information available to investors at the time, he argues that mania era expectations were defensible ex ante: dividends were rising, established lines were profitable, and the collapse required a sequence of adverse developments (capital calls, interest rate rises, revenue disappointments) that investors could reasonably have assigned low probability.[^24] The lesson is uncomfortable for both bulls and bears. Investment in a transformative technology can be individually rational and collectively ruinous at the same time, because the aggregate capacity implied by all the individually reasonable projects exceeds any plausible demand path. No irrationality premium is required for a wipeout, which means the absence of visible madness in today's AI financing is weak evidence of safety. What ruined the railway investor was arithmetic: too many funded miles against too few passengers per year of the demand curve's actual, healthy growth.[^21]

### 2.4 Electricity: the delay between installation and payoff

Electrification contributes a different lesson: even when the technology succeeds and the investment is sound, the payoff arrives with a lag long enough to bankrupt expectations timed to the installation phase. Paul David's classic study of the dynamo observed that electric motors delivered their measured productivity revolution roughly four decades after the technology became available, because capturing the gains required factories to be physically and organizationally rebuilt: the single central steam shaft replaced by distributed unit drive, plant layouts redesigned around workflow rather than around power transmission, and a generation of managers retrained.[^25] Productivity statistics showed little effect through the 1890s and 1900s, then surged in the 1920s when the reorganization completed.

Bresnahan and Trajtenberg's theory of **general purpose technologies** (technologies characterized by pervasive applicability, improvement over time, and the ability to spawn complementary innovations) generalizes the point: a general purpose technology creates most of its value through innovational complementarities in downstream application sectors, whose development follows its own timeline, independent of the core technology's readiness.[^26] For the AI debate this cuts in two directions at once. It licenses patience with today's gap between AI capability and measured enterprise productivity, since the gap is the historical norm for technologies of this class. It simultaneously warns the infrastructure financier that revenue timed to the complementarities' arrival cannot be pulled forward by capex, however large: building the dynamos faster did not make the 1920s arrive in 1900. Capacity installed today against revenue that materializes on the complementarity timeline is exposed for exactly the interval between the two, and Section 11 will argue that interval risk, applied to assets with short useful lives, is the central vulnerability of the current buildout. The dynamo case also fixes the reference class for today's AI productivity skeptics: measured productivity statistics that fail to register a general purpose technology's effect during its installation decades were the norm for electricity, and Section 7 develops the parallel in full through the oil lamp argument.[^25]

### 2.5 Dot-com: an 80 percent drawdown inside a two decade adoption curve

The dot-com episode supplies the modern calibration of how far prices can fall while adoption accelerates. From 1995 to the peak of March 10, 2000, the technology stock basket rose roughly 800 percent; from the peak it fell more than 80 percent.[^27] The list of destroyed vehicles is familiar: consumer dot-coms with negative unit economics, portals priced on eyeballs, and the equipment and carrier complex examined in the next subsection. What the drawdown did to usage is the instructive half: nothing detectable. Internet adoption, traffic, and commerce compounded through the crash and for the two decades after it, along the sequence of broadband rollouts from 2005, the smartphone from 2007, streaming video, cloud computing, and pandemic era remote work.[^27] An investor who in 2002 had concluded from the NASDAQ chart that the internet was overhyped would have been reading the financial layer and calling it the technology, which is precisely the compression error Section 1 predicted for a future AI infrastructure crash. The dot-com case also seeds the base rate for equity selection inside a correct technology thesis: the technology winning is compatible with more than 80 percent of the invested capital losing, because the technology's surplus routes to users and to a small set of survivors rather than to the average share certificate.[^27]

### 2.6 Telecom fiber: the closest analog, quantified

The telecom crash of 2001 to 2002 deserves the most space because it matches the AI infrastructure situation on the most dimensions: a physical capacity buildout, financed by debt and vendor credit, justified by an exponential demand narrative, sold to investors by the infrastructure's own suppliers. Roughly two trillion dollars was spent building 80 to 90 million miles of fiber optic cable in the second half of the 1990s. By 2001, 95 percent of that fiber was dark: installed, buried, and carrying nothing.[^27]

The demand narrative that justified the buildout was specific and false. The industry claim, repeated by executives, analysts, and regulators, held that internet traffic doubled every 100 days. Odlyzko's measurements showed backbone traffic doubling roughly once per year: exceptional growth by any historical standard, and an order of magnitude short of the claim the capacity plans were built on. The overbuild proceeded, in his phrase, in willful disregard of real demand.[^15] When the gap became undeniable, the financial structure imploded: carriers including WorldCom collapsed into bankruptcy through 2001 and 2002, and the equipment suppliers whose revenues had been inflated by vendor financed orders lost nearly everything.[^28] Corning fell from over 100 dollars per share to about 2; JDS Uniphase from roughly 150 to 2.[^27]

Then the deployment phase arrived on schedule. The dark fiber, worthless to its original financiers, was acquired at cents on the dollar by buyers with patient capital, lit progressively over the following decade, and became the physical foundation of broadband, cloud computing, and streaming video.[^29] Traffic never stopped doubling roughly annually through the entire collapse, which is the empirical proof that a demand curve and the solvency of the capacity built for it are separate variables.[^30] Commentary in 2026 has begun mapping this precedent explicitly onto AI data centers, and the mapping is apt with one large caveat, taken up in 2.8.[^31]

### 2.7 The scoreboard

The four infrastructure episodes compress into a table.

| Episode | What investors lost | What the technology delivered afterward | Durable residue of the overbuild |
| --- | --- | --- | --- |
| Railway mania, 1840s | About 50 percent share price decline 1846 to 1850, plus calls on partly paid shares[^23] | A national rail network operated for a century | Track, rights of way, stations |
| Electrification, 1890s to 1920s | Returns deferred for roughly four decades while factories reorganized[^25] | Factory productivity revolution of the 1920s | Generating capacity, grids, reorganized factories |
| Dot-com, 1995 to 2002 | More than 80 percent drawdown from the March 2000 peak[^27] | Two decades of compounding internet commerce and services | Surviving platforms, standards, user habits |
| Telecom fiber, 1996 to 2002 | On the order of 2 trillion dollars; equipment leaders down 98 percent[^27] | Broadband, cloud, streaming on the same glass | 80 to 90 million miles of fiber, lit over 20 years[^29] |

Tulip mania is absent from the table because it left no infrastructure column to fill, which is the reason it belongs in a different genre of financial history.[^22] Across the four that qualify, the constant is the direction of the transfer: the frenzy's capital paid for capacity, the crash repriced the capacity to its economic value, and users plus post crash acquirers captured the difference. The variable is the residue: what physically survives to power the deployment phase.

### 2.8 The honest disanalogy: this cycle's residue is different

The AI application of the template has one structural weakness, and stating it precisely matters more than either ignoring or overweighting it. Rails, dynamos, and fiber were durable: fiber laid in 1999 carries traffic today, and track laid in 1848 carried trains in 1948. The capital good at the center of the AI buildout is the **graphics processing unit** (**GPU**, the parallel processor class on which modern model training and much inference runs), and its useful life is contested. The depreciation debate surfaced publicly in late 2025, when investor Michael Burry argued that hyperscalers depreciating GPU fleets over six years were flattering earnings against a realistic useful life of 2 to 3 years, given the cadence of successive hardware generations.[^32] The counterpoint from operators is that older accelerators cascade into inference and fine tuning service for years after they leave frontier training work, extending economic life beyond the frontier replacement cycle.[^32]

Dave Friedman's analysis of GPU obsolescence draws the conclusion that matters for the Perez template: if a crash comes, the reusable legacy will consist of the shells, the power contracts, the cooling plant, and the grid interconnection rights, while the silicon inside is a fast depreciating consumable.[^33] The deployment phase would therefore run on the durable 60 to 70 percent of the asset stack (site, power, connectivity) refilled with whatever compute hardware the post transition paradigm favors, rather than on the crash discounted chips themselves. We grade the applicability of the Perez renaissance mechanism to AI at B rather than A for exactly this reason: the mechanism is documented twice at full scale, and the asset this cycle overbuilds is only partly durable, so the post crash bargain will be measured in megawatts and permitted sites rather than in petaflops (evidence grade: B).[^33]

### 2.9 What the pattern predicts for 2026 and after

Reading the current cycle against the template produces three dated observations rather than a mood. The frenzy marker is in place: capital expenditure by the largest data center firms is near 750 billion dollars for 2026, up from roughly 450 billion in 2025.[^34] A turning point marker appeared in July 2026: the Philadelphia Semiconductor Index down 28.6 percent from June 22, the Morgan Stanley Momentum TMT Index down 53.5 percent in the month, and the sector's flagship leveraged fund, Situational Awareness LP, forced to liquidate holdings to Citadel on July 30 after falling from 45 billion dollars toward 10 billion.[^35] The boom continuity marker held through the stress: Google's token throughput reached 3.2 quadrillion per month in May 2026, and OpenRouter's routed volume grew 5 fold in six months to 25 trillion tokens per week, with its peak week of May 18 to 24 reaching 28.9 trillion.[^5]

That configuration, financial stress at the infrastructure and momentum periphery against uninterrupted compounding in usage, is the Perez turning point signature as exhibited in 1847 and 2001. The template's prediction is falsifiable and this paper adopts it as such: financial vehicles priced on installation phase assumptions get destroyed or repriced; the physical residue (sites, power, interconnection) changes hands at a discount; deployment then accelerates on cheaper capacity, and the technology's usage curve shows no inflection attributable to the crash. Two developments would falsify the reading. If token consumption growth were to decelerate materially alongside the financial stress, the correction would be a demand event and the aggregate bubble framing would be vindicated. If the current centralized buildout were to remain fully absorbed, with preleased capacity and contracted cash flows carrying projects through their debt schedules as argued by the infrastructure bulls at KKR, the crash leg would simply fail to arrive and installation would hand off to deployment without a break in ownership.[^36] The evidence through May 2026, weighed layer by layer in the next section, is consistent with the template and inconsistent with both falsifiers so far.

## 3. The AI ecosystem as a stack of layers

### 3.1 Reading the stack

The decomposition defended in Section 1 becomes operational here. We divide the AI economy into eight layers: foundation models, semiconductors, cloud platforms, data centers, energy, applications, wrappers, and distribution. For each layer we anchor its current size in post May 2026 figures, identify who holds the risk, and grade how a correction would propagate through it. Two caveats govern the exercise. Layer boundaries are strategic and mobile: the laboratories are integrating upward into applications and downward into compute contracts precisely to escape commoditization of the model layer, so a static map understates the border wars.[^37] And circular financing partially recouples layers whose operating economics differ, meaning financial correlation across the stack is higher than fundamental correlation.[^10] Neither caveat rescues the aggregate label; both are reasons to redraw the map frequently rather than to discard it.

### 3.2 Foundation models

The model layer is where the intelligence is produced and where revenue concentration is starkest. Mid 2026 syntheses of the enterprise stack place OpenAI at roughly 20 billion dollars of annual recurring revenue (**ARR**, the annualized value of ongoing subscription and usage contracts) and Anthropic above roughly 4 billion, figures best treated as order of magnitude given the private status of both.[^38] Two structural facts define the layer's risk profile. First, the majority of OpenAI's revenue has long come from ChatGPT the product rather than from the API (**API**, application programming interface, the programmatic access channel), which means the leading laboratory is economically a consumer subscription company with a research division attached, and its strategy of absorbing application functionality (documented in Section 4) follows from that revenue mix.[^37] Second, training has become a cost center with diminishing returns while inference monetization lags, in the Man Group's February 2026 formulation, so the layer's profitability depends on serving costs falling faster than prices.[^16]

The layer's exposure to a data center correction is asymmetric and, on balance, favorable. Laboratories rent or co-own compute; a capacity glut after a crash lowers their largest cost line. Their genuine risks are competitive (open weight models compressing prices) and financial (dependence on circular equity and compute arrangements with the hyperscalers, which a correction would stress).[^10] Valuation risk at this layer is real and quantifiable against the utility hypothesis, a computation deferred to Section 8; what the layer lacks is the fixed asset stranding that defines the infrastructure layers below.

### 3.3 Semiconductors

Nvidia holds roughly 80 to 90 percent of the AI training chip market, making the semiconductor layer the most concentrated in the stack.[^38] It is also the layer the July 2026 correction hit first and hardest in public markets: the Philadelphia Semiconductor Index fell 28.6 percent from its June 22 peak in under six weeks.[^3] The layer's fundamental question is the durability of the merchant GPU's dominance across both of AI's workloads. Training rewards the GPU's flexibility; inference at scale rewards the efficiency of the **ASIC** (application specific integrated circuit, silicon fixed to one workload), and the 2026 hardware landscape already features Google's TPU v7 generation (**TPU**, tensor processing unit, Google's in house accelerator family) competing against Nvidia's Rubin platform as the workload mix shifts from training toward inference.[^39] The depreciation controversy compounds the paradigm question: if realistic GPU useful life is 2 to 3 years against 6 year accounting schedules, the layer's customers are overstating earnings and understating the true cost of compute, which in turn overstates the sustainable demand for chips.[^32] The semiconductor layer prices growth and paradigm continuity simultaneously; Section 10 argues those are separable bets and the market currently refuses to separate them.

### 3.4 Cloud platforms

The four hyperscalers (Amazon, Google, Meta, Microsoft) plan combined 2026 capital expenditure of approximately 725 billion dollars, up 77 percent year over year: roughly 200 billion at Amazon, 175 to 185 billion at Google, 115 to 135 billion at Meta, and 110 to 120 billion at Microsoft.[^9] That program consumes an estimated 94 percent of their operating cash flow across 2026 and 2027, and their combined free cash flow is forecast to shrink 43 percent between late 2024 and early 2026.[^10] The market has begun charging for the compression: Alphabet's 7 percent single day decline on its July 28, 2026 capex disclosure, with contagion to its three peers, repriced the group's spending as a risk rather than an option.[^2]

The layer's saving characteristics are diversification and position. Hyperscalers operate across several layers at once (silicon, cloud, models, applications, distribution), so a correction in one layer is absorbable by profits from the others; this multi layer redundancy is why our evidence base classifies them among the structural winners of a decomposed correction even while their equities fall with the aggregate trade.[^1] Their genuine vulnerability is the opportunity cost of capital: 725 billion dollars of annual capex directed at a paradigm that shifts underneath it would convert the industry's strongest balance sheets into the largest holders of mispriced capacity, which is why the hyperscalers' own custom silicon programs (each a hedge against the paradigm they are simultaneously buying) are the most informative capital allocation signal in the stack.[^39]

### 3.5 Data centers

This layer is the paper's candidate for the genuine bubble, and the mid 2026 evidence is already split between boom and strain. On the strain side: 30 to 50 percent of American data centers planned for 2026 are expected to be delayed or canceled; of 16 gigawatts (**GW**, billions of watts of power capacity) announced, only about 5 GW are actually under construction; and 150 to 200 billion dollars of data center capex is slipping from 2026 into 2027 and 2028.[^40] The Harvard Belfer Center's February 2026 grid study documents the collision between buildout plans and electric grid capacity, mapping stranded asset and utility risk channels in institutional detail.[^41] On the boom side, the objection is serious and current: vacancy sits near record lows, capacity is preleased years ahead by investment grade tenants, and hyperscalers report being unable to build fast enough for demand.[^42]

Both sides are consistent with the same underlying diagnosis, because bubble risk in physical infrastructure is a vintage and structure question rather than an occupancy question. A facility leased today to an investment grade tenant at today's compute economics is sound; a facility entering a five year construction and interconnection pipeline, financed with debt collateralized by GPUs whose useful life may be shorter than the loan, underwritten at current token prices, is a bet that the entire current paradigm survives its own lead time.[^32] The layer's risk concentrates in exactly the vehicles that history flags: single tenant facilities, **neoclouds** (specialist GPU rental operators financed against their chip fleets), and special purpose developments whose returns require 2030 era demand to arrive in the shape 2026 projected. Wang and Chen's econometrics locate the stack's concerning bubble signals precisely here, where capex acceleration most exceeds monetization.[^1]

### 3.6 Energy

Power has become the binding physical constraint and, with it, the scarce asset of the cycle. Grid interconnection queues in the United States run beyond five years, longer than the planning horizon of most of the compute the connections are meant to serve.[^41] The strain is international: in Ireland, by late 2025, roughly 5.8 billion euros of data center projects were stranded with land and permits secured and no grid connection available.[^43] Utilities face a mirror image risk, documented by the trade press through 2026: generation and transmission built against hyperscaler demand projections becomes a stranded cost on ratepayers if the projected load fails to materialize.[^44]

The energy layer's distinctive property is that its assets survive any AI paradigm. A megawatt of secured, interconnected power is valuable whether the load it serves is a GPU cluster, an ASIC farm, or a non AI workload; substations and interconnection rights carry none of the GPU's obsolescence risk. This is why Section 2's residue analysis matters commercially now rather than only retrospectively: the durable fraction of the current buildout is being assembled at frenzy prices, and the investors who own power, sites, and interconnection through a correction hold the inputs of the deployment phase. The threatened parties in this layer are those financing generation against a single tenant class's projections; the winners are holders of fungible megawatts.[^44]

### 3.7 Applications

The application layer covers products with genuine workflow ownership, proprietary data, or distribution: the stratum above wrappers and below platforms. Mid 2026 investment flows have rotated toward it, on the argument that application and agent companies show nearer term revenue and clearer paths to profitability than infrastructure; Perplexity's valuation around 9 billion dollars with rapid user and advertising revenue growth is the cycle's exhibit.[^45] The rotation is rational on the numbers this paper has already cited: an industry spending roughly 400 billion dollars a year to earn 50 to 60 billion needs revenue bearing layers, and applications are where usage converts to invoices.[^8]

The layer's risk is temporal rather than existential. An application whose margin depends on the current gap between model capability and user need holds an asset with a capability half life measured in months, while its business plan amortizes over years; each frontier release reprices the gap.[^37] The applications that survive successive releases share defensible assets that models do not absorb: regulated workflow ownership, vertical distribution and trust, proprietary data pipelines, and multi model orchestration. One popular defense deserves a discount before it is underwritten: the venture literature's examination of data moats concluded that most claimed data network effects are in practice data scale effects with sharply diminishing returns, so a proprietary dataset defends an application only when marginal data keeps improving the product for the marginal user, a test most corpora fail.[^46] The layer therefore merits underwriting company by company against those specific assets, and the sorting mechanism that performs this filtering continuously, release by release, is the subject of the next subsection and of Section 4.

### 3.8 Wrappers

**Wrappers** (thin applications whose functionality reduces to prompt logic and interface on top of a rented model) are the one layer of the stack this paper concedes to the bubble diagnosis without reservation, and the layer's own dynamics contain the correction. The theory is platform **envelopment**: a platform provider enters an adjacent market by bundling the complement's functionality with its own, exploiting shared user relationships, as formalized by Eisenmann, Parker, and Van Alstyne in 2010 and practiced by platform vendors since Apple made "sherlocked" a term of art two decades ago.[^47] The 2025 to 2026 release sequence supplies dated instances: ChatGPT's Company Knowledge feature absorbed knowledge management wrappers, Claude Code absorbed a band of coding tools, and Claude Tag launched in June 2026 as an embedded enterprise digital worker, each release converting a funded product category into a platform feature.[^37]

Three properties make this layer's excess self liquidating rather than systemic (evidence grade: B for the cadence claim, which lacks a formal mortality dataset).[^48] The purge is continuous, arriving with each release in Schumpeter's perennial gale pattern rather than accumulating toward one crash.[^48] The purged capital is small per event: wrapper businesses are equity light, and their failure transfers users to the platform rather than destroying them. And the wrapper cohort performs a discovery function for the platforms: user innovation research finds lead users pioneered 54.4 percent of the innovations later judged most important, with producers commercializing once the market's extent became clear, and wrapper builders are lead users running priced demand experiments on the laboratories' own telemetry.[^49] The economic identity of the median wrapper founder follows: paying for tokens, cloud, and lab owned coding tools while bearing the market risk, the founder is functionally a premium customer of the platforms, in Srnicek's platform capitalism sense, recruited by survivorship bias that experimental economics has shown produces excess entry even among subjects who know most entrants lose.[^50] The recruiting arithmetic was quantified early: in June 2024, Sequoia's David Cahn computed that capital expenditure at then current levels already implied a 600 billion dollar annual end user revenue requirement, while the largest AI company earned 3.4 billion dollars annualized and most AI startups sat below 100 million; two years later, industry spend near 400 billion dollars against 50 to 60 billion of revenue shows the same structure at larger scale.[^51] The bias operates through the sample the entrants observe. Failures are unobserved, so the visible cohort of successes (a Perplexity or a Cursor class story) systematically overstates the base rate, a distortion the management literature has documented as the peril of benchmarking on survivors,[^52] and one the venture power law makes maximal, since a handful of outliers drive nearly all returns and are therefore the least representative firms in the population.[^53] Sections 4 and 5 quantify this economy; here it suffices to place it on the map as the layer whose bubble deflates on a schedule set by model releases.[^54]

### 3.9 Distribution

The distribution layer decides where intelligence reaches users and where tokens are generated, and it carries the lowest correction risk on the map. Its incumbents own the surfaces: Microsoft through productivity software, Google through Android and the browser, Apple through the device fleet positioned for on device inference at the edge.[^1] Its independent tier is growing at the fastest verified rates in the stack: OpenRouter, which routes demand across models and serving providers, reached 25 trillion tokens per week by May 26, 2026 (peak week 28.9 trillion), 12 times its year earlier volume, and raised a 113 million dollar Series B led by CapitalG on those figures.[^105b] Beneath it, the edge serving infrastructure is compounding: edge inference platforms of the Cloudflare Workers AI class serve dozens of models from hundreds of city locations at sub 100 millisecond latency, and the edge AI market is projected to grow at 21.7 percent annually toward 118 billion dollars by 2033.[^55]

The layer's risk profile is distinctive because it is long usage and neutral on paradigm. A router earns on every token wherever it is generated; an edge network earns more as inference migrates outward from centralized clusters. The serious counterargument holds that large frontier models and batch workloads keep favoring centralized clusters for reasons of latency, model size, power, and data gravity, which caps the migration.[^56] Even under that cap, distribution assets hedge the infrastructure bet rather than compounding it, which is why this paper's synthesis (Section 13) treats routing and edge serving as the sleeper positions of the stack.

The distribution layer is also where the demand side of the paper's thesis becomes measurable. The Jevons paradox, formalized in the ecological economics literature, holds that efficiency gains lower the effective cost of a resource and thereby raise its total consumption; Jevons's original statement from 1866 reads, "It is a confusion of ideas to suppose that the economical use of fuel is equivalent to diminished consumption. The very contrary is the truth."[^57] Microsoft's chief executive invoked the paradox for AI in January 2025, and Erik Brynjolfsson has argued that AI enhanced occupations meet the paradox's conditions.[^57] The token record is the paradox operating in the data: as serving costs fell, consumption rose 330 fold at Google in two years.[^5] The overlooked half of the mechanism is locational. Consumption migrates to whatever serving form is cheapest, and nothing in the paradox routes the induced demand to the incumbent infrastructure; the telecom precedent showed usage doubling annually while the owners of the centralized capacity built for that usage went bankrupt.[^30] Applied forward, this is the paper's claim 8: tokens can grow another 100 fold while the marginal token is generated on an ASIC, an edge network, or a device, leaving late centralized GPU projects without the demand their financing assumed (evidence grade: B, since mid 2026 centralized demand still exceeds deliverable supply and the conjunction is historical rather than yet observed in AI).[^56]

### 3.10 The risk gradient

The map compresses into a gradient, ordered from the layer most exposed to a paradigm correction to the least.

| Layer | Mid 2026 size anchor | Principal risk | Exposure to an infrastructure correction |
| --- | --- | --- | --- |
| Data centers | 150 to 200 billion dollars of capex slipping into 2027 and 2028; 30 to 50 percent of planned US sites delayed[^40] | Lead time versus hardware cadence; single tenant debt structures | Direct: this is the correction |
| Energy | Interconnection queues beyond 5 years; 5.8 billion euros stranded in Ireland alone[^41] | Demand projections failing utilities | High on new generation; low on secured megawatts |
| Semiconductors | 80 to 90 percent training share at Nvidia; SOX down 28.6 percent from June 22, 2026[^38] | Inference migrating to ASICs and edge | High, via order books and multiple compression |
| Cloud platforms | 725 billion dollars of 2026 capex; 94 percent of operating cash flow consumed[^9] | Capital misallocation at paradigm shift | Moderate: multi layer profits absorb the hit |
| Foundation models | About 20 billion dollars ARR at OpenAI; above 4 billion at Anthropic[^38] | Valuation versus utility benchmarks; circular financing stress | Moderate, and costs fall in a glut |
| Wrappers | Purged release by release; 400 billion dollars of industry spend versus 50 to 60 billion of revenue overall[^8] | Envelopment by each model release | Already continuously realized |
| Applications | Rotation of 2026 investment toward the layer; Perplexity near 9 billion dollars[^45] | Capability half life versus business plan duration | Low to moderate, asset dependent |
| Distribution | 25 trillion tokens per week routed; edge AI toward 118 billion dollars by 2033[^105b] | Centralization ceiling on edge migration | Low, and partially inverse |

Read vertically, the table is the paper's thesis in miniature. The financial risk stacks up in the layers that own long lived commitments to a specific hardware paradigm; the usage, and with it the durable economics, stacks up in the layers that touch demand. The remaining sections argue each row in depth, beginning with the layer whose churn is most visible and most misunderstood: the wrapper economy.

[^105b]: Business Wire, OpenRouter announces 113 million dollar Series B led by CapitalG, May 26, 2026, https://www.businesswire.com/news/home/20260526953416/en/

## 4. The wrapper economy and the continuous purge

Of the layers separated in Section 3, the thinnest stratum of the application layer is the one where the bubble diagnosis comes closest to being right. Call it the **wrapper** layer: the population of products whose substance consists of a user interface, a body of prompt logic, and a billing page, assembled on top of a rented foundation model **API** (application programming interface). The thesis of this paper assigns the wrapper layer a specific fate. The layer is genuinely overbuilt, and its excess drains through a mechanism with no counterpart in the historical episodes of Section 2: a continuous, scheduled deflation in which each frontier model release absorbs the weakest cohort of applications built against the previous release. Because the purge arrives in installments synchronized to the laboratories' shipping cadence, the excess never accumulates into the single dramatic collapse that the phrase "the AI bubble bursting" presupposes. This section proceeds from the published theory that predicts the absorption, through the strategic logic that makes absorption deliberate and the documented record of absorptions through 2025 and 2026, to the cadence of the purge, its contrast with crypto's discrete crashes, and finally the economic function the doomed cohort performs while it lives.

### 4.1 Envelopment: the theory that predicted the purge

The mechanism now operating against AI wrappers was formalized more than a decade before the first ChatGPT plugin existed. In their Harvard Business School working paper on **platform envelopment**, later published in the Strategic Management Journal, Eisenmann, Parker, and Van Alstyne described how a platform provider enters an adjacent market by folding the adjacent product's functionality into its own bundle: "Through envelopment, a provider in one platform market can enter another platform market, combining its own functionality with the target's in a multi-platform bundle that leverages shared user relationships."[^58] The theory specifies the conditions under which the maneuver succeeds. Envelopment pays when the platform and the target share a user base, when the platform can replicate the target's functionality at low marginal cost, and when bundling forecloses the target's access to the users both serve.[^58] The enveloped complement loses its distribution before it loses its product, which is why the theory reads as a diagnosis of the wrapper economy written in advance.

Software history supplied the vocabulary before AI supplied the scale. Developers have called the phenomenon being "sherlocked" since roughly 2002 to 2005, when Apple shipped successive versions of its Sherlock search tool that replicated the functionality of third party utilities built for the same operating system, obviating the originals with a bundled release.[^59] The term survives because the pattern recurs on every platform that hosts complements: the operating system absorbs the utility, the browser absorbs the extension, the office suite absorbs the add in. What distinguishes the AI case from the precedent is the intensity of each condition in the envelopment model. Three structural aggravations operate in AI that operated only weakly for Apple and its contemporaries.

First, the laboratory is the wrapper's sole supplier. A sherlocked Macintosh utility at least owned its own code; an AI wrapper's core capability is generated on the platform's servers, purchased per call, and repriced at the platform's discretion. The complement does not merely share users with the platform; it resells the platform.

Second, the platform observes the complement's demand through its own metering. Every wrapper's usage flows through the laboratory's API logs and billing systems, which means the platform holds a real time census of which use cases retain users, at what volume, and at what implied willingness to pay. Envelopment theory already emphasized that platforms attack complement markets where demand is demonstrated and observable through shared users;[^58] in AI, the observation is not inferential but literal, denominated in tokens per account per day.

Third, the cost of bundling falls with every release. Apple had to build each Sherlock feature by hand. A frontier laboratory improves the underlying engine, and capabilities that previously required a wrapper's scaffolding (retrieval logic, tool orchestration, memory management) become native model behavior. The envelopment is partly a by-product of the platform's own research roadmap, which means the wrapper layer erodes even in categories no product manager has explicitly targeted.

### 4.2 Commoditizing the complement

Absorption is strategy as well as side effect. Classical platform economics holds that a platform's value rises when its complements are abundant and cheap, which gives the platform owner a standing incentive to commoditize whatever layer sits between itself and the end customer. In 2026 the laboratories' own revenue mix explains why they act on that incentive from above rather than below. Narayanan and Kapur, in their July 2026 analysis of the laboratories' movement up the stack, document that the consumer product ChatGPT, rather than the API business, has long accounted for the majority of OpenAI's revenue, and that the major laboratories are deliberately constructing switching costs at the application level: enterprise features, memory, organizational knowledge, embedded agents.[^60] The economic logic is straightforward. Raw model access is a commodity exposed to price competition from rivals and open weight alternatives; owning the user relationship at the application layer is the escape from that commodity trap.[^60] The territory into which the laboratories must expand to execute this escape is precisely the territory the wrapper economy occupies. A laboratory that stopped at the API would be building the commodity layer of someone else's business; a laboratory that moves up the stack necessarily displaces its own customers. Narayanan and Kapur's framing carries a warning for buyers as well as builders: the same climb that absorbs wrappers constructs enterprise lock in, since organizational knowledge, memory, and embedded agents raise the cost of leaving the platform in a way raw API access never did.[^60]

This dual position, supplier below and competitor above, is what makes the wrapper layer's risk profile unlike that of any other layer in the Section 3 decomposition. A data center bears technology risk and financing risk; a wrapper bears those plus the risk that its sole supplier absorbs its product as a feature. The supplier can do so at near zero customer acquisition cost, because the users the wrapper laboriously assembled are already, by construction, users of the platform.

### 4.3 The absorption record, 2025 and 2026

The theory would matter little without documented instances. The 2025 to 2026 release record supplies them, with names and dates. The table below lists the absorptions documented by Narayanan and Kapur as of July 2026.[^60]

| First party release | Documented timing | Wrapper category displaced |
|---|---|---|
| ChatGPT Company Knowledge | Documented July 2026 | Knowledge management and "chat with your documents" products |
| Claude Code expansion | Through 2025 and 2026 | Standalone coding assistants and developer copilots |
| Claude Tag, an embedded enterprise digital worker | June 2026 | Enterprise task automation and digital worker products |
| First party memory and agent capabilities | Through 2025 and 2026 | Single feature copilots and personal assistant products |

Each row follows the same sequence. A category of third party products demonstrates demand on the laboratory's API; the laboratory ships the category as a feature; the standalone products compete thereafter against a free or bundled equivalent produced by their own supplier. The products most exposed are the ones the evidence base groups under a single description: any product whose whole value proposition fits inside one future model release note.[^60] Summarizers, document chat interfaces, thin copilots, and prompt orchestration layers all match that description.

The record also supports a nuance the strong form of the claim misses. Moats along this layer form a spectrum rather than a binary. The 2026 investment cycle, as mapped by ValueAddVC's landscape survey, shows capital rotating toward application and agent companies precisely because they exhibit nearer term revenue and clearer paths to profitability than infrastructure plays, with Perplexity valued at roughly 9 billion dollars or more by mid 2026 on rapid user and advertising revenue growth.[^61] Some application companies built on rented models demonstrably capture value. The discriminating variable is what the company owns besides prompt logic. Four asset classes recur among the survivors: proprietary data pipelines that the model cannot regenerate; ownership of regulated workflows where the deliverable is accountability rather than text; vertical distribution and trust in domains such as legal, medical, and insurance; and multi model orchestration that keeps functioning across any single laboratory's upgrade cycle.[^61] A company holding one or more of these assets is an application business exposed to competition. A company holding none of them is a wrapper exposed to absorption. The purge described in this section applies to the second population, which is large.

### 4.4 Deflation, release by release

The temporal signature of this purge is its most consequential property for anyone pricing the "AI bubble". Schumpeter supplied the shape eighty years in advance: creative destruction operates as a continuous process, "the perennial gale", rather than as an episodic storm.[^62] Envelopment theory reinforces the point from the platform's side, since it models envelopment as a repeatable entry strategy the platform can run against successive complement markets, one after another, implying a rolling elimination of complements rather than a one shot event.[^58] The 2025 to 2026 record instantiates both: Company Knowledge, the Claude Code expansion, and Claude Tag arrived as separate, dated releases, each retiring a different category of standalone product, with no single discontinuity anywhere in the sequence.[^60]

For capital allocators the cadence mismatch is the operative fact. Wrapper equity is typically underwritten on venture horizons of five years or more, while the capability releases that reset the layer's competitive map arrived at intervals of months throughout 2025 and 2026.[^60] An investor holding a five year thesis on a product whose margin depends on the current gap between model capability and user need is short the release calendar of the very laboratories the product depends on. The companies structured to survive this exposure are the ones that treat each model gap as a temporary arbitrage: extract cash quickly, hold the option to reposition at every release, and account for the current product as inventory rather than as a franchise.

The same cadence that destroys one organizational form selects for another. The teams that persist across multiple release cycles are those structured for serial product turnover: studios and venture builders that treat each capability gap as a dated arbitrage with a known expiry, price their products to recover capital inside a single release window, and hold engineering capacity ready to redeploy against whatever gap the next release opens. For these operators the purge is an operating rhythm rather than a terminal event, and their existence confirms the deflation thesis from the inside, since a layer in which the rational strategy is capital recovery within one release cycle is a layer whose own participants no longer underwrite durability.

Honesty about the evidence requires one plain statement. No verified quantitative series of wrapper startup mortality per model release exists as of August 2026; the mortality dataset that would let one plot deaths against release dates has not been assembled by any research group we could locate, and the claim of a release synchronized purge therefore rests on the documented mechanism and the documented absorption cases rather than on a measured time series (evidence grade: B). The absence is itself informative about the mechanism. Wrapper deaths are quiet. They take the form of churned subscriptions, abandoned repositories, unpublished markdowns, and small acquihires; they involve no leveraged balance sheets, no bankruptcy filings of systemic size, and no court documents. A purge that proceeds through silent exits generates no series for anyone to compile, which is exactly what the continuous deflation hypothesis predicts and what distinguishes it from the episode examined next.

### 4.5 The crypto contrast

The most recent speculative cycle before AI purged its excess in the opposite tempo, and the comparison isolates what is structurally different about wrappers. Crypto's value destruction clustered into discrete, dated collapse events attached to failing financial vehicles. In May 2022 the Terra stablecoin broke its dollar peg on May 9; Terraform Labs halted the blockchain on May 13; and the collapse erased almost 45 billion dollars of market capitalization in one week, with the LUNA token falling from an all time high of 119.51 dollars to virtually zero.[^63] Six months later, on November 11, 2022, FTX, FTX US, Alameda Research, and more than 100 affiliates filed for bankruptcy in Delaware, following 6 billion dollars of withdrawals in 72 hours, with the exchange's customer shortfall reported at up to 8 billion dollars.[^64]

The difference in tempo follows from the difference in what held the value. Crypto value was stored in leveraged financial claims: exchange balances, algorithmic pegs, interlocked lending books. Claims of that kind fail in a binary way, and their failure transmits through runs and counterparty exposure, which is why the losses printed as single, headline events. Wrapper value sits in the equity of small private companies with little leverage and no interlocking balance sheets. When a release makes a wrapper redundant, the loss is real, but it is local: one capitalization table, one write down, no counterparty, no run. The aggregate consequence is a purge that deflates in installments and never produces the crash print that anchors bubble narratives.

The structural point extends to who bears the losses. Crypto's collapse events socialized losses across retail depositors and token holders, which is what made them systemic and political;[^64] wrapper losses fall on venture equity, the asset class explicitly structured by its power law economics to absorb a high failure rate, as Section 5.3 develops.[^65] A purge that flows through the one asset class designed for it produces neither runs nor rescues, which is a further reason it stays invisible in aggregate statistics.

| Dimension | Crypto purge, 2022 | Wrapper purge, 2024 to 2026 |
|---|---|---|
| Trigger | Failure of financial vehicles (Terra, May 2022; FTX, November 2022)[^63],[^64] | Frontier model and product releases[^60] |
| Tempo | Discrete collapse events dated to the week | Continuous, synchronized to release cadence |
| Transmission | Runs and interlocked balance sheets | Local obsolescence, one category at a time |
| Visibility | Bankruptcy filings, quantified losses | Churn, acquihires, unpublished markdowns |
| Aggregate signature | Single crash prints | Rolling deflation without a single print |

One qualification deserves its full weight, because it comes from the strongest institutional critic of the cycle. Coverage of the Bank for International Settlements (**BIS**, the central bank cooperation body in Basel) systemic risk study in July 2026 argues that the AI boom's circular financing structure could convert the continuous purge into a discrete one: if one large node in the vendor financing cluster slows its purchases, revenue falls across the whole cluster simultaneously, and the funding environment for application startups could seize in one synchronized event.[^66] The point is well taken, and it relocates rather than refutes the argument of this section. The technology cadence produces continuous deflation; the residual risk of a single dramatic purge comes from the financing layer, which is the subject of Section 11. A reader tracking the wrapper layer for systemic signals should watch credit conditions among the large compute buyers, and the release calendars only for the ordinary rolling attrition.

### 4.6 Wrappers as market research for the laboratories

While it lives, the wrapper cohort performs a function that explains why the laboratories tolerate, and indeed cultivate, a layer they will later absorb. The theoretical frame is von Hippel's user innovation research: across 1,678 innovations studied in that literature, lead users pioneered 54.4 percent of those judged most important, with producers commercializing the innovations only "after the extent of the market became clear."[^67] Producers, in this framework, systematically harvest demand discovery performed at users' expense. Wrapper builders are lead users in exactly this sense. They probe thousands of candidate use cases with their own capital, run pricing experiments on their own customers, and absorb the churn of the failures.

The AI version of the pattern carries an aggravation the classical literature did not anticipate: the discovery happens on the platform's own telemetry. A laboratory does not need to commission market research to learn which wrapper categories retain users, because the usage flows through its API metering in real time. Envelopment theory holds that platforms enter complement markets precisely where demand has already been demonstrated and can be observed through shared users;[^58] the 2026 absorption record conforms, since Company Knowledge, the Claude Code expansion, and Claude Tag each shipped into categories first proven at scale by third party products running on the laboratories' own APIs.[^60] Each wrapper cohort thus functions, from the laboratory's side of the meter, as a ranked list of monetizable use cases with pricing experiments and churn data attached.

Two limits on this claim should be stated. The first comes from Casado and Lauten's critique of data moats: most so called data network effects are in practice data scale effects with sharply diminishing returns.[^68] If marginal usage data plateaus in value, then the wrapper cohort's contribution to the laboratories is chiefly the market map, the demonstration of which categories monetize, rather than the raw data itself; and by the same argument, the data assets wrappers hope to defend themselves with are worth less than their pitch decks assume.[^68] The second limit is evidential: no public disclosure documents a laboratory using wrapper API telemetry in its roadmap decisions, so the feedback loop, while structurally implied and behaviorally consistent, is unverifiable from outside (evidence grade: B).

The implications run in both directions along the meter. For founders, a high visibility, single feature product running entirely through one laboratory's API operates as an unpaid product discovery unit for that laboratory, training its own replacement in proportion to its traction. For allocators, traction that is legible to the platform should be discounted rather than celebrated, and the durable positions are those the platform either cannot observe (off platform telemetry, self hosted or open weight deployments) or cannot absorb (regulated distribution, proprietary vertical corpora, customer owned evaluation sets).[^68] The wrapper economy, in short, is a bubble that pays rent to its landlord while inflating, and the landlord collects twice: once on the tokens, once on the market map.

## 5. Founders as customers

The wrapper economy has a second face, visible only when one asks who pays whom while the purge runs. Section 4 described the destruction of wrapper equity from above, by absorption. This section describes the extraction of wrapper cash from below, by metering, and argues that a large share of AI founders occupy an economic position better described as premium customers of the AI platforms than as entrepreneurs capturing AI value (evidence grade: B; the structural revenue asymmetry is documented, while the characterization of the founder population is a supported interpretation rather than a measured statistic). The distinction matters for allocation because it changes what a boom in startup formation measures. If founders are customers, then a surge of AI startups is, in the first instance, a surge in platform revenue, and only contingently a surge in application layer value creation.

### 5.1 Platform capitalism, applied to AI entrepreneurship

The theoretical frame predates the current cycle. In Platform Capitalism, Srnicek describes how platforms position themselves as infrastructural intermediaries and extract value from the activity of the businesses and users who depend on them; the dependent businesses carry the market risk of finding demand, while the platform books recurring revenue from the dependence itself.[^69] The framework was written for ride hailing and social media, and the AI stack matches it with unusual precision because the dependence is metered on four separate lines. An AI startup in 2026 typically pays its platforms through four channels: model API usage priced per token; subscriptions to consumer and team products for its own staff; compute, in the form of **GPU** (graphics processing unit) rentals and cloud capacity; and AI coding tools used to build the product itself. The fourth line has quietly become a laboratory revenue stream in its own right, since Claude Code's explosive growth as a laboratory owned coding product means that even the tooling budget of the builders flows back to the platforms they build on.[^60]

Every one of these lines is invoiced before the startup earns its first unit of revenue. The platform monetizes the attempt; the founder monetizes only the success. Srnicek's framework treats this asymmetry as designed positioning rather than accident: the intermediary structures the market so that participation itself, win or lose, generates platform revenue.[^69] The meter prices are also set unilaterally: API rates, subscription tiers, and compute prices are adjusted at the platforms' discretion, so the startup's gross margin is a residual of decisions taken by its supplier and probable future competitor, a dependence Srnicek's framework describes as the ordinary condition of businesses built on platform infrastructure.[^69] Nothing in this arrangement is hidden or improper, and the same description fits cloud computing's relationship to the startups of the 2010s. What is new in AI is the breadth of the meter (four lines instead of one) and the fact, established in Section 4, that the same counterparty collecting the meter is also the most probable future competitor of the business being metered.

### 5.2 Where the money flows

The aggregate figures show the direction of flow. As a dated baseline: in June 2024, Cahn calculated at Sequoia that each year of AI capital expenditure at then current levels implied an annual end user revenue requirement of roughly 600 billion dollars, at a moment when OpenAI led all AI companies at approximately 3.4 billion dollars of annualized revenue and most AI startups earned under 100 million dollars.[^70] Two years later the gap had widened in absolute terms. Analyses of the 2026 financing structure put AI industry spending at roughly 400 billion dollars per year against 50 to 60 billion dollars of AI revenue, with circular funding arrangements keeping a substantial share of that revenue inside the ecosystem itself.[^71] The revenue that does exist concentrates at the model layer: 2026 syntheses of the enterprise AI stack place OpenAI at approximately 20 billion dollars of **ARR** (annual recurring revenue) and Anthropic above 4 billion dollars.[^72]

Bloomberg's 2026 mapping of the circular deals among Microsoft, OpenAI, Nvidia, and their peers completes the picture: AI revenue largely circulates among the platform and infrastructure players, who alternately invest in, purchase from, and supply one another, while application builders sit outside the loop as net payers into it.[^73] The startup layer, in this flow diagram, is a revenue source for the loop rather than a participant in it.

The meter also runs across the whole life cycle of the attempt. A founding team pays for coding tools during prototyping, for API tokens during product development, for cloud and GPU capacity during scaling, and for production inference for as long as the product lives; each stage of the startup journey has a corresponding platform invoice, and one of the fastest growing laboratory product lines of 2026, Claude Code, is precisely the one that meters the building stage itself.[^60] The historical analogue is the cloud bill of the 2010s startup, and the AI version starts earlier in the company's life and reaches deeper into its cost base, because the tools that write the software, the models that run it, and the infrastructure that serves it can all be invoiced by the same counterparty. A precise measurement one would want, the share of the median AI startup's budget consumed by API, cloud, and tooling payments to the platforms, does not exist in any verified public dataset as of August 2026; the position of startups in the flow structure is documented, while that specific ratio remains unmeasured, and this paper does not estimate it.

The asymmetry between the two sides of the meter can nevertheless be stated without that ratio. The platform's revenue from a cohort of entrants is close to deterministic: every attempt burns tokens, rents compute, and subscribes to tools, so the cohort's aggregate spend arrives whether or not any member finds product market fit. The median entrant's revenue is a draw from a distribution examined in the next subsection, and it is heavily skewed. The platforms hold the picks and shovels position of the gold rush metaphor, formalized into subscription agreements and usage based billing.[^71] One further circularity compounds the pattern on the capital side: venture portfolios exist in which the portfolio companies' largest expense line is payments to platforms in which the same funds are also investors, so the limited partners' capital cycles from the fund through the startup's operating budget into the platform's revenue, and reappears as appreciation of the fund's platform position.[^71]

### 5.3 Survivorship bias and the recruiting function of rare successes

Why does the cohort keep entering, given the economics just described? The experimental literature answered the question before AI existed. Camerer and Lovallo demonstrated in a controlled setting that entrants systematically overestimate their own relative skill and enter markets in excess of profitable capacity, even under conditions where participants know that most entrants must lose money; overconfidence about one's own rank, rather than ignorance of the odds, drives the excess entry.[^74] The result transfers directly: an AI founder does not need to believe the wrapper layer is safe, only that their own product sits above the absorption line, and Camerer and Lovallo's subjects made precisely that structure of error at predictable rates.[^74]

The information environment amplifies the bias. Denrell's analysis of selection effects in management learning explains the mechanism: "Managers learn by example. They, and the consultants they pay for advice, study the methods and tactics of successful companies in search of the magic formulas for business prosperity," while the failed users of the same tactics never enter the sample.[^75] In venture markets the distortion is maximal by construction. Mallaby's history of the venture industry organizes seven decades of evidence around the power law: a handful of outliers drive nearly all returns, which guarantees that the visible successes are extreme tail events and therefore the least representative data points available about the base rate.[^65] The AI builder who studies Cursor, Perplexity, or Midjourney is benchmarking against the far tail of a distribution whose body is invisible.

What converts this ordinary bias into a recruiting function is the platforms' direct interest in the inflow. In earlier cycles, a platform benefited diffusely from founder enthusiasm; in this one, every recruited builder is incremental API, cloud, and tooling revenue from the day they start building, independent of their eventual outcome.[^71] Success stories circulate through developer marketing, model demonstration events, and accelerator programs, and each circulation recruits the next cohort of metered customers.

The timing interaction with Section 4 sharpens the cost of the bias. Accelerator cohorts assemble around the visible winners of the last release cycle, which means they clone the categories whose demand is already proven, legible to the platform through its own telemetry, and therefore next in line for absorption; the cohort is recruited into precisely the position the envelopment mechanism targets, at the moment its window is closing.[^60] The information asset that would correct the bias, an honest base rate, is also the one nobody currently sells at scale. No verified 2026 dataset of AI startup failure rates exists, and the analytical products that would compound in credibility if the purge becomes visible (base rate datasets, post mortem repositories, layer aware market maps) remain largely unbuilt. The bias mechanism is established at evidence grade A; its specific magnitude among 2026 AI builders is not, because no verified 2026 dataset of AI startup failure rates exists, and the calibration of current overconfidence therefore rests on the structural argument rather than on measured attrition (evidence grade: C).

The steelman deserves its space. ValueAddVC's 2026 landscape reports the investment cycle rotating toward application and agent companies precisely because they show nearer term revenue and clearer paths to profitability than earlier cohorts, across a mapped population of more than 200 companies and over 150 billion dollars raised, with value concentrating in a few names.[^61] If application layer base rates have genuinely improved, then some fraction of 2026 builder confidence is rational updating on better odds rather than bias. The honest synthesis is that both claims hold simultaneously: the odds may have improved at the top of the application layer, per the rotation ValueAddVC documents,[^61] while the visible sample from which the median entrant learns remains a tail selected sample, per Denrell and Mallaby,[^75],[^65] so the population of entrants still overshoots whatever the true base rate now is.

### 5.4 The retail investor analogy

The structure described above has a familiar shape from financial markets, and this paper offers it explicitly as an analogy rather than an identity (evidence grade: C). The AI founder economy reproduces the architecture of retail brokerage. A platform sells market access to a mass of small participants and earns on their activity volume rather than on their outcomes; a small, highly visible cohort of winners is circulated as marketing material; the recruitment of each new cohort is funded, in effect, by the losses and fees of the previous one; and the median participant underperforms while the platform's revenue compounds with participation itself. Substitute token metering for commissions, laboratory demonstration events for brokerage advertising, and accelerator cohorts for the inflow of new accounts, and the mapping is complete. The Camerer and Lovallo result even supplies the same behavioral engine in both markets: participants enter on overconfidence in their relative rank, in full knowledge of the aggregate odds.[^74]

The analogy has honest limits, and stating them strengthens rather than weakens it. A founder who fails exits with skills, shipped products, and domain knowledge, assets with durable value that a losing trade does not leave behind. Entry can therefore be positive sum for the ecosystem even where it is negative in expectation for the median entrant, since the failures collectively perform the use case discovery documented in Section 4.6 and train the labor force of the next cycle. And the 2026 rotation evidence shows a real cohort graduating into revenue generating businesses rather than remaining metered hopefuls.[^61]

The analogy also identifies the structural counter position. In retail brokerage, the participants who escape the platform's economics are those who stop paying for access; in AI, the equivalent is the builder who moves lines off the platform meter. Operators of open weight models (the Llama, Mistral, and DeepSeek lineages), local inference stacks, inference cost optimizers, and hybrid deployments that keep data and spend in house all reduce the platform's take from the cost base, and they simultaneously escape the telemetry legibility that Section 4.6 identified as the absorption trigger, since usage the platform cannot observe is usage it cannot map. The laboratories' own escape from the commodity trap, documented in the Up the Stack analysis, implies the mirror strategy for builders: consume the commodity layer where it is cheapest, and keep the differentiated layer (data, workflow, and distribution) off the platform's meter.[^60]

What the analogy earns, in exchange for its limits, is a set of testable predictions and allocation rules. It predicts that founder participation tracks the visibility of winners with a lag, and that a visible purge event, should the financing shock of Section 4.5 materialize, would contract entry sharply rather than gradually.[^66] It implies that platform revenue from the builder segment is procyclical to narrative, and should be underwritten as such rather than as a stable annuity. For allocators it yields three operating rules, delivered in order. First, read aggregate AI startup formation partly as a platform revenue indicator, since the cohort's spend arrives at the platforms regardless of the cohort's fate.[^71] Second, examine venture portfolios for the circularity in which portfolio companies' dominant expense is platforms the same fund also holds, because that structure books the same capital twice.[^71] Third, underwrite application companies against the survival asset list of Section 4.3 (proprietary data, regulated workflows, vertical distribution, multi model orchestration) rather than against traction metrics that are fully legible, and therefore fully actionable, by the platform that meters them.[^61]

## 6. The demand for intelligence as an expanding frontier

Sections 4 and 5 described how value drains within the application layer's thinnest stratum and who collects it on the way down. Nothing in that account implies weakness in the demand for intelligence itself, and this section argues the opposite: the demand for intelligence behaves as a moving frontier rather than a fixed market of existing tasks, so every valuation exercise that sizes AI against its current use cases systematically understates the future market. The argument runs from the observed failure of fixed market forecasts, through the price and volume evidence that identifies the demand structure, through the economic theory of task creation, to the behavioral mechanism by which each solved capability raises human ambition toward the next, with web development serving as the completed case study.

### 6.1 The fixed market fallacy

Every bearish aggregate on AI embeds a demand model, usually implicit: enumerate today's use cases, price them, sum them, and compare the sum to the capital deployed. The model has a measurable track record in this cycle, and the record is one of systematic, rapid undershoot. In May 2026, Goldman Sachs Research estimated global token consumption at 5.6 quadrillion tokens per month; by July 2026, tracking of disclosed throughput put the global run rate at roughly 11 quadrillion tokens per month, double the estimate within two months.[^76] An institutional forecast of the frontier's position was stale before a quarter had passed. The same tracking shows the mechanism at the level of individual firms: Fireworks AI, an inference provider, processed about 15 trillion tokens per day in April 2026 and about 40 trillion per day by mid July 2026.[^76] Whatever the correct model of intelligence demand is, a fixed total addressable market (**TAM**, the standard sizing concept) is demonstrably the wrong one: planning built on a fixed TAM in this market fails within quarters rather than years.[^76] The pattern generalizes beyond one bank's estimate: across the 2026 record, institutional forecasts of AI usage have been revised upward within months of publication, a track record examined further in Section 6.6.[^76]

### 6.2 Jevons in the token data

The demand structure that does fit the data is the one William Stanley Jevons described for coal and that the efficiency literature calls the rebound effect: the **Jevons paradox**, in which efficiency gains in the use of a resource raise rather than lower its total consumption, because the induced demand outweighs the savings. A 2026 SSRN (Social Science Research Network) working paper applies the framework directly to AI infrastructure and measures the elasticity: a 1 percent decrease in the price of compute generates a 1.42 percent increase in volume demanded.[^77] Demand for machine intelligence is, on this measurement, super elastic: price declines expand total spending rather than shrinking it.

The 2025 to 2026 price and volume record instantiates the signature. Blended token prices fell approximately 67 percent year over year, from 18.40 dollars to 6.07 dollars per million tokens between the first quarter of 2025 and the first quarter of 2026, while volumes exploded across every disclosed meter.[^78] The volume side of the ledger, assembled from post May 2026 disclosures, reads as follows.

| Meter | Earlier reading | Later reading | Multiple |
|---|---|---|---|
| Google, tokens processed per month | 9.7 trillion (May 2024) | over 3.2 quadrillion (May 2026) | 330x in two years, 7x year over year[^79] |
| Global run rate, tokens per month | 5.6 quadrillion (Goldman Sachs estimate, May 2026) | roughly 11 quadrillion (mid 2026) | 2x in about two months[^76] |
| Fireworks AI, tokens per day | 15 trillion (April 2026) | about 40 trillion (mid July 2026) | 2.7x in three months[^76] |
| Goldman Sachs projection, tokens per month | 5 quadrillion class base (2026) | 120 quadrillion (2030 projection) | 24x, driven by agentic inference[^80] |

A second SSRN paper by the same author closes the loop on spending: despite cost declines of one to two orders of magnitude over four years, total enterprise AI spending keeps growing, because deployment volume expands to absorb every efficiency gain, and enterprise budgets restructure from labor line items toward compute line items as they do so.[^81] This is the precise, measured refutation of the argument that cheaper intelligence deflates the intelligence market. Cheaper intelligence has so far meant more spending on intelligence, at an elasticity now estimated above unity,[^77] and the implications of that same elasticity for the infrastructure layer, where it is often misread as an unconditional guarantee, are taken up in Section 12.

The distribution of the volume growth identifies who collects it. Google alone processed over 3.2 quadrillion tokens per month by May 2026, a 330x multiple in two years, which places a hyperscaler with first party models at the head of the meter;[^79] the laboratories' token businesses expand with every use case the frontier absorbs; Nvidia holds the toll position on the compute that serves the volume; and specialized inference providers ride the same curve, as Fireworks AI's near tripling of daily throughput in three months shows.[^76] The demand expansion of this section therefore feeds the revenue concentration documented in Section 5.2: super elastic token demand is collected through a small number of meters, which is one reason expanding aggregate demand coexists with a purging application layer.

### 6.3 Task automation and task creation

Price elasticity alone cannot carry the frontier thesis, because elasticity measured on existing tasks could in principle saturate once those tasks are fully served. The load bearing question is whether the set of economically valuable tasks is itself fixed. The canonical treatment answers that it is endogenous. Acemoglu and Restrepo's task based framework models two opposing forces: automation displaces labor from existing tasks, while a reinstatement effect creates new tasks in which labor and new technology are complementary, so the task set expands as technology advances rather than standing still to be consumed.[^82] Their own empirical decomposition adds the honest caveat: over the three decades preceding 2019, displacement accelerated in the United States while reinstatement weakened, so the expansion of the task set is a documented tendency rather than an automatic law.[^82] In the AI instance the reinstatement channel is directly observable in the labor market for tasks that did not exist in 2022: agent orchestration, model evaluation, and the verification work described in Section 6.5 are new task categories in Acemoglu and Restrepo's sense, complements created by the automation itself.[^82]

For AI specifically, the task creation channel has a second storey the general framework lacks. Cockburn, Henderson, and Stern argue that deep learning is a general purpose technology of a particular kind: a general purpose "invention of a method of invention" (**IMI**), a technology that reshapes the discovery process itself, and that AI "may have an even larger impact by serving as a new general purpose method of invention that can reshape the nature of the innovation process and the organization of R&D."[^83] An IMI does not merely automate existing research tasks; it creates research tasks that had no human equivalent, and the market for intelligence then includes the output of every field the method transforms. The compounding version of this argument is formalized in a 2025 National Bureau of Economic Research (**NBER**) model of feedback loops in innovation networks: in a calibrated semi endogenous growth simulation, full automation of software research combined with just 5 percent automation elsewhere generates explosive growth within six years, because automated research productivity in one sector raises research productivity in the others.[^84] That result is a model calibration rather than an observation, and this paper treats it as an upper envelope, not a forecast.

The serious counterweight is Acemoglu's own macro estimate: modeling the exposed task base and the quality of feasible automation, he puts the total factor productivity (**TFP**, output growth unexplained by measured inputs) gain from AI at no more than 0.66 percent cumulatively over ten years, later revised below 0.53 percent, on the grounds that most exposed tasks are hard to automate well and genuinely new task creation is slow.[^85] The estimate deserves engagement rather than dismissal, and two replies carry the weight. The first is that the near term forecasting contest is already being scored: the token record of Section 6.1 shows realized demand doubling institutional estimates within two months,[^76] and the field experiments of Section 6.4 show workers attempting new work rather than merely doing old work cheaper, both of which sit awkwardly with a model in which the task frontier barely moves. The second is that a modest TFP path and an enormous intelligence market are mutually consistent: TFP measures economy wide productivity, while the market for intelligence can grow by restructuring existing expenditure, which is precisely what the enterprise spend data shows, with budgets shifting from labor to compute at constant or growing totals.[^81] Acemoglu's number, even taken at face value, bounds the macro dividend; it does not bound the revenue pool available to the sellers of intelligence.

### 6.4 From A to A+1 to B: the ambition mechanism

The thesis statement of this paper compresses the demand dynamic into a sequence: solving task A creates demand for task A+1, and the completion of a task family shifts ambition to family B. The field experimental literature supplies the measured core of that sequence. In the largest of the experiments, 758 consultants at Boston Consulting Group were randomized into AI assisted and control conditions across realistic tasks. Inside the frontier of the model's competence, consultants with GPT-4 access completed 12.2 percent more tasks, worked 22 to 28 percent faster, and produced results at 38 to 43 percent higher quality; and the study's central construct, the jagged frontier, expands with each model generation, resetting the set of tasks workers attempt at all.[^86] The same study measures the boundary condition that keeps the mechanism honest: when workers pushed delegation beyond the model's actual capability, performance dropped by an average of 19 percentage points, so ambition raised faster than capability produces errors rather than output, and the frontier dynamic is conditional on capability genuinely advancing.[^86]

Two further experiments complete the pattern at both ends of the skill distribution. Noy and Zhang found professional writing time fell 40 percent with quality up 18 percent, and, more telling for the ambition mechanism, exposure to the tool changed what participants chose to produce afterward.[^87] Cui and coauthors, running field experiments on software developers at Microsoft, Accenture, and a Fortune 100 firm, measured a 26.08 percent increase in completed tasks, with the strongest uptake among less experienced developers.[^88] In all three settings, capability gains translated into more attempted and completed work rather than into the same output produced with less effort.

The behavioral record outside the laboratory shows users climbing the sequence in real time. ChatGPT approached 1 billion weekly active users (**WAU**) in July 2026, from roughly 400 million eighteen months earlier, and passed 1 billion monthly active users in June 2026.[^89],[^90] On July 9, 2026, OpenAI launched ChatGPT Work, an agent that carries a goal through to finished documents and sites and stays on complex projects for hours;[^91] the product category exists because users hand better models bigger jobs the moment the frontier moves. Each step up the sequence multiplies the resource intensity of demand: Gartner's March 2026 analysis, as reported in the token demand tracking, finds agentic workloads consuming 5 to 30 times more tokens per task than chatbot interactions.[^76] The mechanism can be stated precisely. Solving task A converts A from a priced deliverable into a cheap input; the priced unit migrates to A+1, the larger job of which A is now a component; and when an entire task family is colonized, ambition relocates to family B, at a token multiple per task documented between 5x and 30x.[^76] Direct measurement of rising aspiration as such does not yet exist; the mechanism is inferred from the adoption and workload escalation data above (evidence grade: B).

The token multiple also locates the layer the frontier rewards next. If agentic workloads consume 5 to 30 times the tokens of chat per task,[^76] then the binding constraints on the next leg of demand are orchestration and evaluation: the tooling that decomposes large goals into agent runs, verifies intermediate output, and decides when a multi hour job has actually succeeded. These layers serve use cases that barely existed before 2026, which positions them on task families created by each step up the ambition sequence rather than competed away by it, the sleeper position of the demand frontier.

### 6.5 Web development as a solved frontier

One task family has already traversed the full sequence, and it deserves examination as the completed experiment: web development, the construction of ordinary sites and applications. By mid 2026 this task class sits substantially inside the frontier. Laboratory owned coding agents carry a specification to a deployed site: ChatGPT Work is marketed on exactly that capability, finished documents and sites from a stated goal, sustained over hours of autonomous work;[^91] Claude Code's growth through 2025 and 2026 made it one of the documented absorption engines of Section 4.[^60] The capability that once defined a professional occupation's daily output is now generated on demand, priced in tokens.

The fixed market model makes a clear prediction for this case: demand for web development, once the task is solved, should collapse toward the cost of tokens, and the market should shrink. The observed sequence ran the other way at every measured point. Completed development work rose where the tools arrived, by 26.08 percent in the multi firm field experiments, with the gains concentrated among less experienced developers, meaning the entry price of building for the web fell and participation widened.[^88] The workloads did not stay at the old task size: coding and agentic workflows are precisely the category Gartner measures at 5 to 30 times the token consumption of chat,[^76] and inference providers serving these workloads posted the throughput multiples of Section 6.2 (Fireworks AI's tokens per day rising 2.7x between April and mid July 2026).[^76] And the priced unit moved up a level: when producing a page or a feature approaches zero marginal cost, the scarce inputs become specification, integration, and verification, the A+1 of which page production is the component. The value of the solved task did not evaporate; it was absorbed into a larger unit of work that consumes strictly more intelligence per engagement. Professional writing traversed a version of the same sequence earlier: once drafting time fell 40 percent at higher quality in the experimental setting, the priced unit moved toward editorial judgment and verification, and exposure to the tool changed what writers subsequently chose to produce.[^87] This reading of the web development record as the template for successive task families is the paper's interpretation of the cited measurements rather than a measured generalization (evidence grade: B), and it is the single most direct piece of evidence against the proposition that capability improvements saturate demand: the most fully solved professional task class in the AI economy is also among its heaviest and fastest growing consumers of tokens.[^76]

### 6.6 Where ambition migrates next, and what the frontier does to market sizing

The theoretical reason to expect the web development pattern to repeat lies in the definition of a **general purpose technology**: Bresnahan and Trajtenberg characterize such technologies by pervasiveness, continuous improvement, and innovational complementarities, and show that a general purpose technology's value is realized in downstream applications that do not exist when the technology first appears, so early use cases mechanically understate the final market.[^92] The historical calibration of the lag comes from David's study of the dynamo: electricity's productivity payoff took roughly four decades to arrive, because factories captured it only after redesigning production around unit drive motors in the 1920s, two generations after the enabling technology.[^93] Both results are structural rather than cyclical, which is why this paper treats them as load bearing despite their age.

Applied to 2026, the framework implies that the intelligence market cannot yet be measured, only bracketed, and the bracket is wide enough to prove the point. Estimates scoped to vendor revenue put the 2026 global AI market at roughly 601 to 750 billion dollars,[^94] while estimates scoped to total AI spending, infrastructure included, reach 2.52 trillion dollars for 2026, up 44 percent year over year,[^95] a fourfold divergence produced entirely by scope choices. Longer horizon projections extend the bracket rather than resolving it: 1.81 trillion dollars for the AI market proper by 2030 in one widely used series,[^96] and approximately 3.5 trillion dollars by 2033 in another.[^97] A market whose measured size swings by 4x on definitional choices, and whose usage denominator doubles against institutional forecasts within two months,[^76] is a market still in the discovery phase of its demand curve. The claim that current use cases understate the future market by an order of magnitude follows from the general purpose technology analogy and the observed forecast dispersion rather than from direct measurement, and carries that grade (evidence grade: B).

Where the next task families lie is partly visible in the research record. Peer reviewed work reported in May 2026 shows AI assistants accelerating scientific discovery by designing and interpreting experiments rather than merely executing them,[^98] and Stanford's Institute for Human Centered AI documents AI transforming discovery across fields in the manner of the telescope and the microscope, an instrument class change in what questions can be asked at all.[^99] These are the early entries of task families with no human baseline, the family B of the sequence, and Section 7 examines them as the petrochemical stage of the current oil lamp era.

The same migration repositions existing industries around the new task families. Research services whose billable unit is the manual experiment cycle, contract research organizations and contract materials laboratories among them, face at the level of physical work the envelopment dynamic that Section 4 documented at the level of software, as experiment design and interpretation move inside the instrument.[^98] On the other side of the migration sit the AI native discovery companies, built on automated hypothesis generation and closed loop laboratories, whose tasks exist only because AI exists; and the research groups working on non text modalities (world models, robotics, biology), which occupy the position David's framework assigns to the unit drive motor: the component whose maturation lets the rest of the economy redesign itself around the new input.[^93] The speed of the migration is genuinely uncertain, and the paper reports the uncertainty as measured: in the International AI Safety Report's 2026 elicitation, AI forecasting experts assign a median 20 percent probability to compressing six years of AI progress into two, while superforecasters assign 8 percent.[^100]

The frontier thesis carries one warning for allocators, and it is the note on which this block closes. An expanding frontier does not pay every position on it. David's dynamo lag ran four decades,[^93] and any AI position priced off current use case revenue, with a runway shorter than the gap between the present wave and the downstream applications that will dominate final value, can be right about the technology and still be carried out of the market before the value arrives; the threatened class here is short duration capital in long duration theses. The demand for intelligence is an expanding frontier at every measured point in the 2026 record.[^76] Which layers of the stack convert that expansion into recoverable returns, and on what timetable, is a separate question, and Sections 7 and 8 take it up where the demand side leaves off.

## 7. The oil lamp stage

### 7.1 What general purpose technology status actually claims, and what it does not

A **general purpose technology** is a technology that satisfies three conditions simultaneously: it spreads across most sectors of the economy, it keeps improving over a long horizon, and it makes innovation cheaper in the sectors that adopt it. Bresnahan and Trajtenberg formalized the category in 1992 and published it in the Journal of Econometrics in 1995, using steam, electricity and semiconductors as the reference cases.[^101] Throughout this paper the phrase is written out rather than abbreviated, because the three letter form collides with the name of a model family.

The part of that paper which matters for capital allocation is not the taxonomy. It is the market failure result. Bresnahan and Trajtenberg show that a general purpose technology and its downstream applications form a system of innovational complementarities: the value of improving the core technology depends on how many application sectors have already invested in adapting it, and the value of adapting it in any one sector depends on how good the core has become. Decentralized markets coordinate that loop badly. Their conclusion is that complementary innovation arrives "too little, too late" relative to the social optimum in the early phase.[^101] The economic significance of the technology therefore shows up with a structural delay that has nothing to do with the quality of the core technology and everything to do with the slow accumulation of downstream adaptation.

This has a direct consequence for how an investor should read a revenue statement in 2026. If the general purpose technology model is correct, the revenue that a core technology generates in its first decade measures the state of downstream adaptation, not the eventual size of the category. A low number is compatible with an enormous endpoint, and the theory says the number should be low.

The objection to placing weight on that structure comes from Acemoglu, who accepts that artificial intelligence belongs to the general purpose technology class and still bounds its measured effect. His task based accounting gives at most a 0.66% gain in United States **total factor productivity** (TFP, the residual of output growth not explained by measured labor and capital inputs) over ten years, revised downward to less than 0.53%.[^102] The structure of the argument matters more than the number: Acemoglu builds up from the share of tasks that current systems can perform and the average cost saving on those tasks. That method is exactly the one the general purpose technology literature says will understate the endpoint, because it prices the technology against today's task inventory. It is also the method that would have produced a small number for electricity in 1900. Acemoglu's estimate is a rigorous lower bound conditional on the task set staying roughly fixed, and it is not evidence that the task set stays fixed.

### 7.2 The dynamo delay

Paul David gave the delay its canonical measurement in 1990. Commercial electric power dates from the early 1880s. By 1900 dynamos were visible, in his phrase, everywhere but in the productivity statistics.[^103] United States manufacturing did not capture the productivity gains of electrification until the 1920s, roughly forty years after the technology became commercially available.

The mechanism David identified is the one that matters here. Early electrified factories did not redesign anything. They removed the steam engine and put a large electric motor in its place, driving the same overhead line shafts, the same belts, the same building layout. The factory was still organized around the geometry of mechanical power transmission, which forced multi story buildings, dense machine placement near the shaft, and a plant layout dictated by the drive train rather than by the flow of work. The gains arrived with unit drive, where each machine received its own motor. Unit drive made the building layout free, allowed single story plants organized around material flow, and eventually enabled the moving assembly line. Capturing electricity's value required rebuilding the physical and organizational form of production, and that rebuild waited for the old capital stock to depreciate and for a generation of factory architects to learn the new form.

The analogue in 2026 is legible. Most enterprise deployment of language models reproduces the existing workflow with a model attached at one step: the model drafts the email that a person was already going to write, summarizes the document that a person was already going to read, generates the code that a person was already going to type. The randomized evidence in Section 8 measures exactly this shape of gain, and it is real and modest. The organizational form in which the technology is native, where the unit of work is a task delegated to a process that runs for hours without supervision, only began shipping as a mass product in July 2026.[^104] David's history says the interval between the line shaft phase and the unit drive phase is where the returns live, and that the interval is measured in decades rather than quarters.

### 7.3 Kerosene before petrochemicals: the 1909 crossover

The petroleum industry supplies the sharper version of the same pattern, because in oil the eventual dominant product was physically present from the beginning and thrown away.

Commercial drilling in the United States begins in 1859. The business it created was illumination. Kerosene remained the predominant commercial end use for petroleum refined in the United States until 1909, when motor fuels exceeded it.[^105] For the first fifty years of the industry, oil was a lamp business. Daniel Yergin's history of the industry documents that era in detail: the competitive battles, the Standard Oil consolidation, and the refining economics were all organized around lamp oil.[^106]

The instructive detail is gasoline. In the early refining era gasoline was an unwanted byproduct of kerosene production, volatile and dangerous, with no market of consequence.[^105] Refiners dumped it. The molecule that would become the largest fraction of the barrel and the foundation of twentieth century transport was a disposal problem for half a century, because the machine that consumed it did not exist yet. The market for gasoline was not discovered by refiners searching for applications of their waste stream. It was created from outside the industry by the internal combustion engine and by the mass produced automobile, and the refiners then reorganized the entire physical plant around cracking heavier fractions into the product they had previously discarded.

Petrochemicals repeat the structure one level further out. Plastics, synthetic fibers, fertilizers and pharmaceutical feedstocks are all built from hydrocarbon streams that the lamp oil industry had no reason to isolate. None of it was visible from inside the illumination business.

The disciplined form of the analogy is this. In 1900, an analyst pricing the oil industry against total lamp oil revenue would have produced a defensible valuation of the industry and a badly wrong valuation of petroleum. The error would not have been in the arithmetic. It would have been in the assumption that the current end use enumerated the future end uses.

### 7.4 Where artificial intelligence actually sits on that curve in mid 2026

The lamp phase claim is testable against usage data rather than asserted. The evidence is that the dominant deployed use of frontier models remains conversational text for personal purposes.

ChatGPT passed one billion **monthly active users** (MAU, distinct accounts using the product in a calendar month) in June 2026, serving roughly 2.5 billion prompts per day.[^107] Usage compilations of OpenAI's own research published in July 2026 put approximately 70% of ChatGPT usage outside work contexts, with roughly 49% of all queries falling into the general knowledge "asking" category.[^108] A product with a billion users, half of whose queries are general knowledge questions asked for non work reasons, is an information utility of considerable value and it is not a production input to the physical economy.

Set that against the capital being formed. Worldwide artificial intelligence spending is forecast at 2.52 trillion dollars for 2026, 44% above 2025, while application layer revenue remains an order of magnitude smaller.[^109] A June 2026 multi method evaluation of whether artificial intelligence is in a financial bubble frames the 2026 configuration explicitly in Perez's terms, as an installation period in which financial capital builds the infrastructure of a technological revolution well ahead of the deployment period in which the revolution's productive value is realized.[^110] Perez's own framework, built on five technological surges across three centuries, describes precisely this gap between capital formation and realized downstream value as the signature of installation.[^111]

The first visible step past the lamp arrived on July 9, 2026, when OpenAI released ChatGPT Work, an agentic product designed to field tasks that run for hours rather than answer questions that resolve in seconds.[^104] One product launch is a data point, not a phase transition. It is the first mass market artifact whose unit of work is a delegated task rather than a conversational turn.

### 7.5 The candidate aviation fuels

If the current use is the lamp, the question for allocation is where the automobile equivalent is being built. The honest answer in 2026 is that the highest quality results outside text are concentrated in science, and that none of them are commercial products yet.

The lineage starts with structure prediction. AlphaFold, published in Nature in 2021, was the first computational method to regularly predict protein structures at atomic accuracy even in the absence of a similar known structure, and was competitive with experimental structures in a majority of CASP14 targets.[^112] CASP is the Critical Assessment of protein Structure Prediction, the blind community benchmark that had defined the field's difficulty for a quarter century. The result mattered because it converted a problem that consumed years of laboratory work per structure into a computation.

The same method class then moved into materials. GNoME, published in Nature in 2023, produced 381,000 new stable materials, almost an order of magnitude more than all prior work combined.[^113] It moved into formal mathematics: AlphaProof reached silver medal standard at the 2024 International Mathematical Olympiad, the first medal level performance by an artificial system, with the method published in Nature in 2026.[^114]

Through 2026 the results shifted from single system demonstrations to closed loop discovery. Two Nature papers reported in May 2026 describe multi agent systems that run the discovery cycle rather than a single step of it.[^115] Google DeepMind's Co-Scientist proposed novel drug candidates and combination therapies for acute myeloid leukemia, identified new liver fibrosis targets, and proposed antimicrobial resistance mechanisms. FutureHouse's Robin identified candidate treatments for age related macular degeneration.[^115]

Stanford's Human Centered Artificial Intelligence institute assembled the broader inventory on May 27, 2026, and the range is the point. The Samudra model simulates 1,000 years of climate per day on a single graphics processor, against 12 years per day for traditional models, a throughput ratio of roughly 83 to 1. The EVO deoxyribonucleic acid language model generated 16 novel bacteriophages that outperformed natural ones against resistant bacteria. A Virtual Lab of artificial agents designed COVID antibody binders superior to prior human designs within days. The largest neuroscience dataset assembled recorded 3.1 million neurons. Deep learning identified 3 to 4 million regulatory deoxyribonucleic acid elements. In mathematics, systems achieved perfect scores on the hardest undergraduate examinations in December 2025 and solved dozens of open Erdos problems.[^116]

Two properties of this list are load bearing. First, every entry is measured in scientific output and none is measured in revenue, which is the signature of an application layer that has not been monetized rather than one that has failed to find a market. Second, none of these results came out of a chat product. The organizations producing them are research units, and the capability being demonstrated is search over a structured space under a verifiable objective, which is a different technical problem from conversational assistance.

Robotics belongs in the same category by argument rather than by published result. A foundation model that controls actuators converts artificial intelligence from an information good into a labor input, which changes the accounting category it occupies in an economy. The evidence base for robotics foundation models in mid 2026 is thinner than the evidence base for scientific discovery, and this paper does not assign it a proof level above D (evidence grade: D).

### 7.6 The objection that deserves the most weight

The strongest counterargument to the entire oil lamp framing is not that the science results are weak. It is that scientific tractability does not convert into economic output at the rate the analogy implies.

Risa Wechsler, quoted in the Stanford review, states the constraint precisely: artificial intelligence changes what problems are tractable, and it does not tell researchers what problems matter.[^116] Downstream of that, the conversion from a computational prediction to an economic good runs through experimental validation, wet lab throughput, regulatory approval, and manufacturing. A model that proposes a thousand candidate molecules does not compress the clinical trial that each candidate must survive. The bottleneck migrates from hypothesis generation to physical verification, and physical verification scales with laboratory capital and calendar time.

Acemoglu's bound is the quantitative expression of the same skepticism.[^102] His argument that the measurable task base is bounded implies that the lamp framing assumes the existence of the automobile equivalent rather than demonstrating it. That assumption is not free. Some technologies did stall at their early phase: supersonic passenger transport and nuclear fission for industrial process heat both had credible expansion paths that never opened.

The correct epistemic position is therefore graded rather than confident. The historical pattern that a general purpose technology's early applications understate its endpoint is strongly documented across steam, electricity and semiconductors (evidence grade: A for the pattern). The existence and quality of high value artificial intelligence results outside text is directly established by peer reviewed publication (evidence grade: A). The claim that those results become the largest economic applications of the technology is a forward projection supported by structure rather than by revenue (evidence grade: B). The specific timing of the crossover is not estimable from available evidence, and this paper does not estimate it.

### 7.7 What the oil lamp reading implies for allocation

Three consequences follow, and they are not symmetric in confidence.

The first concerns companies whose entire market is the current lamp equivalent. A text and chat only artificial intelligence company faces the exposure that lamp oil refiners faced from 1900: competition from a superior substitute for the existing use, combined with the reorganization of the industry around a product the incumbent was not built to make. The exposure is structural rather than imminent, and it is sharpened by the fact that the chat interface is the layer most directly absorbed by each frontier model release.

The second concerns the compute and infrastructure layer. Bresnahan and Trajtenberg's innovational complementarity argument implies that infrastructure with optionality across unknown future applications is worth more than infrastructure specialized to known ones, because the identity of the eventual dominant application is precisely what nobody knows. This is the strongest argument available for owning general capability rather than application specific capacity, and Block 4 examines the conditions under which it fails, namely when the infrastructure's technical form is itself specialized to a training paradigm that changes.

The third concerns the science layer. Companies building artificial intelligence for non text domains are currently priced against scientific output rather than against revenue, because they have little revenue. The set includes Google DeepMind's science units, Isomorphic Labs, FutureHouse in automated biological discovery, genomics regulatory mapping, catalysis and crystallography startups, and climate simulation firms. If the analogy holds even partially, this is where the gasoline is, and it is the part of the stack that a 2026 valuation screen cannot see, because the screen reads revenue.

## 8. Valuations against the utility hypothesis

### 8.1 Stating the condition that the valuations require

The claim that large artificial intelligence company valuations are rational is not a claim about sentiment. It is a conditional statement with a testable antecedent: the valuations clear if artificial intelligence becomes a utility scale spending category, meaning a recurring per person and per employee expenditure that continues in a recession, rather than a discretionary software purchase that gets cut in a procurement review.

Trammell and Korinek supply the formal endpoint of that condition. Their model of economic growth under transformative artificial intelligence, revised in April 2026, examines regimes in which artificial systems automate essentially all work. Automating production alone raises the growth rate substantially and lowers the labor share of output.[^117] That is the only regime in which a software company can plausibly command a utility scale fraction of world output, because it is the regime in which the company's product substitutes for labor rather than assisting it. The model is a description of a regime, not a forecast that the regime arrives.

Damodaran approaches the same question from the valuation side. His June 2026 analysis of SpaceX, OpenAI and Anthropic in the context of index inclusion argues that judging these companies on current earnings metrics misses where young companies' value actually sits, which is in the option value of a large future revenue base rather than in a current multiple.[^118] That is the standard and defensible case for accepting large forward multiples on young companies.

The objection arrives from the same author. Cornell and Damodaran's Big Market Delusion, published in the Financial Analysts Journal in 2020, describes the specific failure mode of a genuinely enormous market.[^119] When a **total addressable market** (TAM, the full revenue opportunity if a product captured every potential buyer) is perceived as vast, every participant gets priced as though it will capture a large share of it. Each individual valuation is defensible against the market size. The sum of the valuations implies a set of market shares that exceeds one hundred percent. The market can be exactly as large as claimed and most of the equity in it still mispriced. This objection is not a rebuttal of the utility hypothesis. It is a rebuttal of the inference from the utility hypothesis to any particular company's price.

### 8.2 The revenue trajectories, with the conflicts stated

The post May 2026 numbers are strong and the sourcing on valuations is inconsistent enough that a serious reader should see the disagreement rather than a reconciled figure.

Anthropic's annualized revenue rose from approximately 9 billion dollars at the end of 2025 to approximately 47 billion dollars by May 2026, a factor of roughly 5.2 in five months.[^120] The company reports gross of cloud reseller payouts, which inflates the figure relative to peers reporting net, and the comparison to any other lab's number should carry that caveat. On valuation the sources conflict directly. Sacra associates a 965 billion dollar post money valuation with Anthropic's June 1, 2026 initial public offering filing.[^120] A separate 2026 compilation lists Anthropic at a 183 billion dollar valuation with 30 billion dollars of annualized recurring revenue as of April 2026.[^121] The two figures are roughly a factor of five apart across a two month interval. This paper does not adjudicate between them. An allocator using either number should verify it against the filing itself, and any argument whose conclusion flips between 183 billion and 965 billion is an argument that has not been made.

OpenAI's figures carry the same structure. Sacra reports a 852 billion dollar post money valuation from the round closed March 31, 2026, more than one billion monthly active application users as of May 2026, enterprise revenue above 40% of the total and on track for parity with consumer by the end of 2026, gross margin of approximately 33% constrained by inference cost, and projected cash burn of 27 billion dollars in 2026 rising to 63 billion in 2027.[^122] A separate compilation gives 730 billion dollars post money and 122 billion dollars raised as of April 2026.[^123] The revenue figure most often quoted, approximately 25 billion dollars of **annualized recurring revenue** (ARR, the most recent period's revenue extrapolated to a full year) as of February 2026, predates this paper's May 2026 freshness cutoff and is used here only as dated history.[^122]

Two of those OpenAI numbers deserve more attention than the valuation. A 33% gross margin is a utility's margin structure, not a software company's, and it is set by inference cost. A projected 63 billion dollar cash burn in 2027 means the company's equity value depends on continued access to primary capital markets through at least that year.

On xAI the freshness rule bites and the evidence base has a hole. The available figures are a 250 billion dollar valuation following the SpaceX transaction of February 2, 2026, a 20 billion dollar Series E closed in January 2026, and approximately 3.83 billion dollars of consolidated 2025 revenue of which about 500 million dollars is standalone excluding X advertising.[^124] All of these predate May 2026 and no verified post cutoff xAI figure was obtainable. The paper therefore does not present a current xAI valuation or revenue run rate. What survives the freshness rule is the structural observation, which does not depend on the exact numbers: xAI is the clearest case in the sector of a valuation whose defense rests entirely on category size rather than on current revenue, which is the configuration Cornell and Damodaran describe.[^119]

The growth rate comparison is the most informative fresh datum in the set. Claude's monthly user growth ran at 640% year over year in the second quarter of 2026 against 62% for ChatGPT.[^122] Two products in the same category growing at rates an order of magnitude apart, at very different absolute bases, indicates a market still in share formation rather than one that has settled into stable positions.

### 8.3 What consumers actually value the product at

The consumer side of the utility question has a proper measurement rather than an inference. Nguyen, Brynjolfsson, Kazinnik, Collis and Eggers ran online choice experiments on representative United States adults in July 2025 and March 2026, measuring **willingness to accept** (WTA, the compensation a person requires to give up a good) for one month without access to artificial intelligence chatbots.[^125]

Mean monthly willingness to accept rose from 98 dollars in July 2025 to 124.50 dollars in March 2026, an increase of 27% in eight months. Aggregate United States consumer surplus from the category rose from 116 billion to 172 billion dollars over the same interval.[^125]

Set the mean against the household bill benchmark. The average United States household mobile phone bill in 2026 is 98 dollars per month, cable and internet 124 dollars, electricity 125 dollars, with total household bills of 3,289 dollars per month.[^126] The mean measured value of chatbot access in March 2026 sits between the mobile phone bill and the electricity bill.

The distribution destroys the naive reading of that comparison. The same study reports a median willingness to accept of 11.40 dollars per month against the mean of 124.50 dollars.[^125] The ratio of mean to median is nearly eleven to one. A minority of intensive users carries almost the entire measured value of the category, and usage frequency is the strongest predictor of an individual's willingness to accept. For the median American adult, artificial intelligence is worth roughly a tenth of a phone bill.

Prices reflect that split. Claude's consumer pricing as of August 2026 runs from free, to Pro at 20 dollars per month (17 dollars on annual commitment), to Max starting at 100 dollars per month for the 5x and 20x usage tiers.[^127] The Max tier crosses the average mobile phone bill. The Pro tier sits at one fifth of it. The observed conversion rate is the missing variable: ChatGPT paid users reached approximately 50 million, roughly 5.5% of the user base, as of the first quarter of 2026, up from 20 million in April 2025.[^128] That figure predates the freshness cutoff and is the most specific conversion measurement available; treat it as the state of the first quarter rather than the state of today.

The gap between a mean measured value of 124.50 dollars and a modal paid price of 20 dollars at a 5.5% conversion rate is where consumer revenue growth is available without any new capability at all. It is also where the pricing power of the labs will be tested, because capturing it requires either raising prices on subscribers who currently pay 20 dollars or converting users who have revealed, by not subscribing, that their value is nearer the 11.40 dollar median.

### 8.4 The telecom benchmark, used correctly

Telecom is the right benchmark for a utility scale consumer category, and the comparison has to be made on the correct axis.

The structural scale of the reference category is documented but dated: global telecommunications services revenue was projected at 2.2 trillion dollars in 2015 rising to 2.4 trillion by 2019.[^129] No verified post May 2026 global telecom average revenue per user figure was obtainable during research for this paper, and this paper does not use one. The argument proceeds on penetration instead, which is measured and current.

Ericsson's June 2026 Mobility Report records that fifth generation mobile subscriptions passed 3 billion worldwide.[^130] Telecom is a near universal bill: essentially every economically active adult on the planet pays one, every month, in every macroeconomic condition, and has done so for decades.

That is the bar, and it is a bar on penetration rather than on price. ChatGPT's one billion monthly active users at 5.5% paid conversion implies on the order of 50 million paying consumers worldwide.[^107],[^128] Mobile telephony has billions. The distance between the two is not a price problem. Artificial intelligence subscription pricing already reaches and exceeds telecom pricing at the top tier.[^127] The distance is a universality problem, and universality is what makes a bill non discretionary.

Two mechanisms could close it, and they differ in how much they require. The first is organic conversion, which requires the median user's value to rise from 11.40 dollars toward the price point, a movement the 2025 to 2026 measurements show beginning but far from complete.[^125] The second is bundling, in which carriers and internet service providers include artificial intelligence access in phone and broadband plans, which converts the category from a discretionary subscription into a line inside the telecom bill itself. Bundling reaches universality without ever winning a purchase decision, and it transfers the margin to the carrier. The doxo data explains why this matters: the average household already commits 3,289 dollars per month to bills, so a new 100 dollar line item is taken from an existing category rather than added to a household's total.[^126]

The honest statement of the consumer case is therefore split by user segment. For intensive users, measured value and available price tiers both already exceed the telecom bill (evidence grade: B). For the median individual, median willingness to accept of 11.40 dollars per month is an order of magnitude below the average phone bill, and the utility framing is an analogy rather than a measurement (evidence grade: C).

### 8.5 Enterprise spend per employee, where the measurement is strongest

The enterprise side is where the utility hypothesis has direct rather than indirect support, because per employee spend is now measured across a large sample.

The Ramp AI Index of June 2026, covering more than 70,000 United States businesses, reports a median company artificial intelligence spend of 2,246 dollars per month and 11.38 dollars per employee per month. The top decile of companies spends 611 dollars per employee per month. The top one percent spends 7,450 dollars per employee per month. Mean company spend was 140,842 dollars per month in April 2026, which against a median of 2,246 dollars identifies a power law distribution. Token consumption across the tracked businesses grew 1,001% from January 2025 to April 2026.[^131]

The spread between the median and the top decile is the most important number in this section. A ratio of 611 to 11.38, roughly 54 to 1, means the category is not a seat license. Seat licenses compress spend across companies because everyone pays the same price per head. Usage based inference does the opposite, and the top of the distribution is already spending per employee what enterprise software historically spent per employee per year.

The driver is identified: agentic coding workloads. Typical artificial intelligence coding usage runs 150 to 250 dollars per developer per month, rising to 500 to 2,000 dollars per month for agentic workflows consuming 400,000 to 2 million tokens per session.[^131] Against that, seat prices understate cost by construction. Microsoft 365 Copilot lists at 21 dollars per user per month standard, with a promotional 18 dollars per user per month on annual commitment through September 30, 2026, 25.20 dollars month to month, and a bundled Business Premium with Copilot at 32 dollars per user per month annual.[^132] Claude Team runs 25 dollars per standard seat monthly (20 dollars annual) and 125 dollars per premium seat (100 dollars annual), with Enterprise from 20 dollars per seat plus usage that scales with model and task.[^127] A 21 dollar seat with 500 dollars of monthly inference behind it is a 521 dollar line item.

The infrastructure side of that consumption is visible at the inference providers. Menlo Ventures reported on July 16, 2026 that daily token volume at one inference provider went from 15 trillion to 43 trillion per day since late 2025, a near tripling, driven by models gaining the ability to sustain long running tasks.[^133] Token volume growth of that shape is a direct measurement of the shift from conversational turns to delegated tasks described in Section 7.

The market structure underneath the spend, from Menlo's enterprise decomposition, gives Anthropic 40% of the enterprise model application programming interface market, OpenAI 27% and Google 21%, with Anthropic at an estimated 54% in coding, which is the highest per seat spend category.[^134] The same source measured enterprise generative artificial intelligence spend of 37 billion dollars in 2025, growing 3.2 times year over year, split 19 billion dollars applications and 18 billion dollars infrastructure, with 47% of artificial intelligence deals reaching production against 25% for traditional software as a service.[^134] Those 2025 figures are dated history under the freshness rule and are used here for the decomposition rather than for the level.

### 8.6 What an enterprise seat actually buys, measured

Enterprise spend is durable only if it buys something measurable. Two randomized studies establish the floor.

Dillon, Jaffe, Peng and Cambon ran a randomized trial of Microsoft 365 Copilot across more than 6,000 workers in 56 organizations over six months. Access produced 30 minutes less time reading email per week and documents completed 12% faster, with approximately 40% regular adoption among workers granted access.[^135] Cui, Demirer, Jaffe, Musolff, Peng and Salz measured three field experiments with software developers and found a 26.08% increase in completed tasks, the highest per seat productivity effect measured in a professional setting.[^136]

The two results bracket the category. General knowledge work gets a real but bounded gain, concentrated in communication overhead, with adoption at 40% of licensed seats. Software development, where the output is verifiable and the task decomposes cleanly, gets a gain roughly an order of magnitude larger in economic terms. That is why coding is the highest spend category and why Anthropic's coding share matters to its revenue mix.

The counterweight is realized firm level return. McKinsey's State of Artificial Intelligence survey found 88% of organizations using artificial intelligence in at least one function, 39% reporting enterprise level earnings before interest and taxes impact, and 5.5% qualifying as high performers.[^121] Adoption is near universal and financial impact is concentrated in a minority. Enterprise per employee spend can therefore rise while the spend is durable only in the firms that captured the impact, which is a composition risk inside a growing aggregate rather than a challenge to the aggregate itself.

### 8.7 Where the revenue gap genuinely bites

The aggregate arithmetic is where an allocator should concentrate. Combined 2026 hyperscaler capital expenditure is running at 700 to 725 billion dollars, examined in detail in Section 9.[^137] Worldwide artificial intelligence spending is forecast at 2.52 trillion dollars for 2026.[^109] Anthropic at approximately 47 billion dollars annualized and OpenAI at a run rate whose most recent verified figure predates the cutoff are the two largest pure model revenue lines in the sector.[^120],[^122]

The gap between capital formation and model layer revenue is roughly an order of magnitude, and the correct reading of that gap is specific rather than general.

It does not indict the category. Section 7 established that a general purpose technology's early revenue measures downstream adaptation rather than eventual size, and the enterprise measurements in this section show a spend line growing 1,001% in fifteen months with a top decile at 611 dollars per employee per month.[^131] A category whose heaviest users already spend at utility scale is not a category without demand.

It bites in three specific places. The first is the median firm, where realized earnings impact is reported by 39% of organizations while 88% deploy, which means a large share of current enterprise spend is unvalidated by financial return and is therefore cuttable.[^121] The second is the marginal lab, which is the Cornell and Damodaran point applied: if two or three platforms capture the utility scale category, the aggregate private valuation of the sector exceeds any consistent view of it, and the loss falls on the companies priced against category size without a defensible share.[^119] The third is the financing structure of the leaders themselves. A 33% gross margin against a projected 63 billion dollar cash burn in 2027 makes OpenAI's equity value a function of capital market access as much as of product demand.[^122]

The conclusion the evidence supports is conditional and specific. Model layer revenue is compounding at rates no enterprise software category has matched, and no post May 2026 data yet shows artificial intelligence spending behaving like a non discretionary utility bill for the median consumer or the median firm (evidence grade: B). The valuations require the top of the distribution to broaden. The observable test is whether median per employee spend rises toward the top decile, and whether median consumer willingness to accept rises toward the mean. Both are measured annually by the sources cited here, which makes the hypothesis falsifiable on a known schedule.

## 9. The scaling paradigm and its limits

### 9.1 Scaling laws as the intellectual foundation of the buildout

The 700 billion dollars of 2026 hyperscaler capital expenditure rests on a specific empirical result, and the result is real.

Kaplan and colleagues at OpenAI published it in January 2020. Language model loss falls as a power law in three inputs taken separately: model size, dataset size, and training compute. The relationships hold across more than seven orders of magnitude, and architectural details matter far less than scale.[^138] The practical content of that finding is that capability became forecastable. An organization could commit capital to a training run and predict, within a usable band, the loss the resulting model would achieve, before the run started. Predictability is what converts research into a capital budget.

Hoffmann and colleagues at DeepMind refined it in 2022 with the result known as Chinchilla. Training over 400 models spanning 70 million to 16 billion parameters and 5 billion to 500 billion tokens, they found that compute optimal training scales parameters and tokens in equal proportion: doubling model size requires doubling training tokens.[^139] The demonstration was decisive. Chinchilla at 70 billion parameters, trained on four times more data, outperformed Gopher at 280 billion parameters and GPT-3 at 175 billion, reaching 67.5% on the Massive Multitask Language Understanding benchmark.[^139] Chinchilla is the canonical proof that data ingestion is a binding input alongside parameter count, and it is the reason the data wall discussed below is a constraint on the paradigm rather than a curiosity.

The measured trajectory that follows from acting on these laws is steep. Amodei and Hernandez measured a 3.4 month doubling time in the compute used by the largest training runs between 2012 and 2018, a growth factor exceeding 300,000 over the period.[^140] Sevilla and colleagues documented the deep learning era compute explosion across three eras of machine learning.[^141] Epoch AI's current measurement is that training compute of frontier models grows 4 to 5 times per year, and that training costs of the largest models double roughly every eight months.[^142],[^143] Epoch projects the number of models trained above 1e26 **floating point operations** (FLOP, the standard unit of computational work) growing from about 10 by 2026 to over 200 by 2030.[^144]

### 9.2 The capital expenditure that prices in the paradigm's permanence

Combined 2026 **capital expenditure** (capex, spending on long lived physical assets) plans of Amazon, Alphabet, Microsoft and Meta run to approximately 700 to 725 billion dollars, up about 77% from roughly 410 billion dollars in 2025. The decomposition is Amazon at approximately 200 billion, Alphabet at 175 to 185 billion, Meta at 125 to 145 billion, and Microsoft at 110 to 120 billion.[^137],[^145]

That capital is not a bet on artificial intelligence demand in the abstract. It is a bet on a specific technical configuration: that capability will continue to be produced by very large training runs on general purpose accelerators in centralized facilities, and that the marginal value of the next order of magnitude of training compute justifies its cost. Every element of that configuration is an empirical claim that could stop being true while demand for the output keeps rising.

The market has begun to test it. Alphabet raised its 2026 capex forecast alongside second quarter earnings in late July 2026, and the shares fell 7% the following day, with Amazon, Meta and Microsoft declining in sympathy.[^137] A capex increase read as a growth signal in 2024 was read as a margin risk in July 2026. The change in the market's interpretation function is itself information about how durable investors believe the scaling return to be.

### 9.3 Brute force described precisely

The phrase "brute force" is used loosely in commentary, so it is worth stating exactly what the current training procedure does.

A frontier pretraining run assembles a corpus on the order of ten trillion tokens or more, drawn from web crawls, digitized books, code repositories, licensed feeds and increasingly from synthetic generation. DeepSeek-V3, whose technical report is unusually specific, trained on 14.8 trillion tokens.[^146] The model is then shown that corpus, in most cases essentially once, and its parameters are adjusted by gradient descent to reduce the error in predicting each next token. No part of the procedure identifies which tokens carry information the model lacks. The corpus is processed in full, at uniform effort per token, and the learning signal per token is whatever the prediction error happens to be.

The inefficiency of this is measurable against the only other system known to acquire language competence. Frank's work in Trends in Cognitive Sciences puts the gap at four to five orders of magnitude: language models receive between ten thousand and one hundred thousand times more language data than human children who reach comparable linguistic competence.[^147] A factor of ten thousand is not a tuning difference. It indicates that the procedure extracts a small fraction of the information each token contains, and compensates with volume.

That is the precise sense in which the paradigm is brute force. It is not that scaling fails to work. Kaplan and Chinchilla show it works with unusual reliability. It is that the conversion rate from data to capability is low, and a low conversion rate is an efficiency target that engineering effort will attack.

### 9.4 The data wall

Volume compensation has an arithmetic limit, because human generated text is a finite stock.

Villalobos, Ho, Sevilla, Besiroglu, Heim and Hobbhahn at Epoch AI estimated the effective stock of quality adjusted public human text at approximately 300 trillion tokens, and projected full utilization between 2026 and 2032 if training trends continue.[^148] The window opens in the year this paper is written. The estimate is a range rather than a date because it depends on how aggressively labs count lower quality sources and on how many epochs they are willing to run over the same data.

The obvious escape is synthetic generation, and it has a documented failure mode. Shumailov, Shumaylov, Zhao, Gal, Papernot and Anderson showed that recursive training on model generated content causes model collapse: the tails of the distribution disappear first, then the model converges toward a narrowed version of its own output, and the degradation is irreversible in their setting.[^149] The failure is structural rather than incidental, because a model's samples systematically under represent the low probability regions of the distribution it was trained on.

The ambient environment makes this sharper. Estimates circulating in 2026 put artificial intelligence generated material at roughly half of new internet content.[^150] If that figure is even approximately right, the distinction between "crawl the web" and "train on synthetic data" is eroding without anyone choosing it, and the premium on verified human data, licensed archives, and contractor produced corpora rises accordingly.

The data wall does not stop capability improvement. It changes which input is scarce. When tokens were abundant and compute was scarce, the optimization problem was how to buy more compute. When tokens are scarce and compute is abundant, the optimization problem becomes how much capability to extract per token, which is a research problem rather than a procurement problem.

### 9.5 Algorithmic efficiency, measured

The rate at which that research problem has been solved historically is the single most important number in this section, because it governs how fast a given capability becomes cheap.

Ho, Besiroglu, Erdil, Owen, Rahman, Guo, Atkinson, Thompson and Sevilla measured it directly using more than 200 language model evaluations spanning 2012 to 2023. The compute required to reach a fixed performance level has halved roughly every eight months, with a 95% confidence interval of 5 to 14 months.[^151] That is faster than the historical doubling cadence of transistor density. The same study decomposes the sources: the transformer architecture alone is worth nearly two years of algorithmic progress, and Chinchilla optimal scaling is worth 8 to 16 months.[^151]

The measurement replicates in vision. Hernandez and Brown found that the compute needed to reach AlexNet level ImageNet performance fell 44 times between 2012 and 2019, a doubling of efficiency every 16 months, of which hardware improvement alone would have delivered only 11 times.[^152] Erdil and Besiroglu extended the methodology across vision benchmarks and corroborated persistent multiplicative software gains.[^153]

The objection is in the same body of work and it is important. Epoch's Shapley decomposition attributes 60 to 95% of observed performance gains to compute and data scaling, against 5 to 40% to algorithms, and notes that many algorithmic innovations are themselves scale dependent, meaning they only pay off at large model sizes.[^154] The correct reading is that efficiency substitutes for compute at a fixed capability level and complements compute at the frontier. Getting today's capability for a tenth of the cost is an efficiency result. Getting tomorrow's capability still requires the large run.

That distinction is exactly the one that matters for infrastructure demand, and Section 10 develops it. An eight month halving at fixed capability means that the compute required to serve the median workload falls by roughly a factor of a thousand per decade, while the compute required to advance the frontier keeps rising.

### 9.6 The price collapse as the observable consequence

The efficiency result shows up in prices, where it is easiest to verify.

Inference for GPT-4 class capability cost approximately 20 to 30 dollars per million tokens in late 2022 and 2023. By spring 2026 the same capability class ran at roughly 0.40 to 0.80 dollars per million tokens, a decline of about 95% in two years and approximately a factor of a thousand over three.[^155] Google's Gemini 3.1 Flash was priced in April 2026 at 0.10 dollars per million input tokens and 0.40 dollars per million output tokens, which analyses put at a 99.7% reduction over three years.[^156]

The 2026 analyses attribute the decline to three compounding forces delivered in this order: algorithmic efficiency in training and serving, better hardware across the H100 to Blackwell transition and custom accelerators, and price competition from open weight models. The composite rate is approximately a factor of ten per year at a fixed benchmark score since 2021.[^157]

A tenfold annual price decline at fixed capability is the central economic fact of the model layer. It means that a data center project financed against the token prices of its planning year faces revenue per unit of compute that falls by an order of magnitude per year unless the workload migrates to higher capability tiers at the same rate. It also means that any application built on a fixed capability tier sees its cost of goods sold collapse, which is the mechanism behind the token volume growth measured at the inference providers in Section 8.

### 9.7 Distillation, compression, and the turn to pattern extraction

The techniques that produce the price collapse are worth naming, because they are the concrete content of "efficiency" and they point at where training is going.

Distillation trains a smaller model to reproduce the behavior of a larger one, transferring capability without transferring parameter count. Quantization reduces the numerical precision of weights and activations, cutting memory and bandwidth requirements. Pruning removes parameters that contribute little to output. **Mixture of experts** (MoE, an architecture in which each token is routed to a small subset of the model's parameters rather than all of them) reduces the compute per token at a given total parameter count. Speculative decoding uses a small model to draft tokens that a large model verifies in parallel. Each of these attacks the same target: the cost of a unit of capability delivered.

The deeper shift is on the training side, and it changes what the data is for. Sorscher, Geirhos, Shekhar, Ganguli and Morcos showed at NeurIPS 2022, in an Outstanding Paper, that with a good data pruning metric the error can fall exponentially rather than as a power law in dataset size.[^158] That result breaks the scaling regime rather than accelerating within it. A power law says that halving the error requires multiplying the data by a large constant factor. An exponential says it does not. The condition is a metric that identifies which examples carry information the model has not extracted, which relocates the difficulty from data acquisition to data selection.

DeepSeek supplied the applied evidence at frontier scale. DeepSeek-V3 is a 671 billion parameter mixture of experts model with 37 billion parameters active per token, trained on 14.8 trillion tokens in 2.788 million H800 graphics processor hours, the run widely quoted at approximately 5.6 million dollars of compute cost.[^146] The engineering, the architecture and the data recipe substituted for raw expenditure at a ratio that repriced assumptions across the industry.

DeepSeek-R1, published in Nature 645 in 2025, went further in the direction that matters for the data wall. Reasoning capability was incentivized by pure reinforcement learning without human labeled reasoning trajectories.[^159] The model generated its own attempts, a verifier scored them, and the training signal came from the score. That procedure manufactures training data from a specification of correctness rather than consuming a finite stock of human text, which is the same structural move that AlphaZero made in games. It applies wherever correctness is checkable, which covers mathematics, code, formal proof and a growing set of scientific tasks, and it does not apply where correctness is a matter of taste.

The 2026 research frontier extends the move. QUEST trains deep research agents using only 8,000 fully synthesized tasks and approaches or surpasses frontier closed agents on research benchmarks.[^160] Eight thousand tasks against fourteen trillion tokens is a difference of nine orders of magnitude in data volume for a specific capability class. The 2026 training data ecosystem is correspondingly described as six parallel sourcing layers, spanning licensed feeds, contractor generated data and synthetic generation, with the center of gravity moving away from raw web crawls.[^161]

### 9.8 The open weight diffusion channel and moat deflation

The efficiency turn would matter less to infrastructure economics if the techniques stayed inside the labs that discovered them. They do not, and the diffusion rate is measured.

Epoch AI's capability index comparison, published May 29, 2026, finds that since January 2026 the most capable open weight models lag frontier closed models by an average of four months, or 8 points on the Epoch Capability Index, a gap comparable to the distance between two consecutive frontier releases from the same lab.[^162] A four month lag on the best freely downloadable model is a four month depreciation clock on any capability lead.

The objection is documented and should be held alongside the headline. A LessWrong analysis argues the public benchmark comparison understates the gap: on private held out benchmarks open models lag 8 to 10 months against 4 to 6 on public ones, which is consistent with open models being tuned toward published evaluations.[^163] Epoch itself measured the lag widening from three months in October 2025 to four months in May 2026.[^162] Commoditization of the capability layer is well evidenced and the frontier premium has not vanished.

The speed of technique diffusion is separately measured, and it is faster than the capability lag. Zhang and colleagues surveyed the replication of DeepSeek-R1 style reasoning approximately 100 days after its release, covering supervised fine tuning recipes and reinforcement learning from verifiable rewards reproduced by the open community.[^164] Hugging Face launched Open-R1, a fully open reproduction pipeline covering training, evaluation and synthetic data generation, within days of the release.[^165] TinyZero reproduced emergent reasoning behaviors in models as small as 1.5 billion parameters for under 30 dollars of compute.[^166] A frontier training technique moved from proprietary result to undergraduate exercise inside one quarter.

The distribution channel is large and its usage skews small. Hugging Face reported more than 2.5 million models and over 700,000 datasets as of May 2026, with the company reaching profitability.[^167] Earlier 2026 figures put the hub at over 2 million public models with roughly 40,000 new models landing per month.[^168] Models under 1 billion parameters account for approximately 92% of hub downloads, around 2.2 billion cumulative.[^169] Actual deployment concentrates in models small enough to run outside a hyperscaler data center.

Grogan's March 2026 working paper draws the strategic conclusion: pretraining at scale is not a durable competitive moat, because open alternatives approach frontier performance, inference costs collapse, application companies consume models as commodities, and open weights serve state sovereignty objectives that guarantee continued funding for them.[^170] Commentary through 2026 consolidates the same reading of open weights as the commoditizing force in the stack.[^171] The paper's claim about full erosion of infrastructure moats is an extrapolation that the frontier lag data only partly supports (evidence grade: B).

The complementary question is where advantage goes if it leaves capital. Zegart and Johnston's analysis of DeepSeek's team found 31 top contributors averaging over 1,500 citations each, approximately 98% trained in Chinese institutions, and concluded that technological leadership will be won with strategies for global talent competition alongside faster chips and bigger models.[^172] The organization is roughly 150 researchers and engineers in flat, goal based groups without rigid hierarchy.[^173] Against that, the 2026 talent market shows capital buying skill rather than skill displacing capital: Meta reportedly offered a single researcher up to 1.5 billion dollars over six years in June 2026, with 100 million dollar signing bonuses documented for a small set of senior hires, and packages reaching 300 million dollars over four years for more than a dozen researchers hired in January 2026.[^174] Compensation surveys across the major labs report equity at 40 to 70% of pay.[^175] When a research bench is priced like a company, the largest balance sheets acquire it, and know how concentrates with capital rather than dispersing away from it.

### 9.9 Is the paradigm transitional, and what would that do to training demand

The signals that the current paradigm is transitional converge from independent directions and none of them is decisive alone.

The data stock is finite and its exhaustion window opened in 2026.[^148] The sample efficiency gap against human learners is four to five orders of magnitude.[^147] Practitioner syntheses in 2026 report capability specific plateaus, with knowledge tasks flattening beyond roughly 30 billion parameters and code beyond roughly 34 billion, alongside frontier training cost projections of 5 to 10 billion dollars per run.[^176] Hooker's 2026 thesis paper documents cracks in the scaling relationships and argues the next competitive wave is defined by the cost of adaptability rather than by model size, citing Falcon 180B from 2023 being outperformed by Llama 3 8B one year later, a twenty fold reduction in parameters at higher capability.[^177] Alternative architectures exist and work: Mamba, a selective state space model with no attention mechanism, achieves five times higher inference throughput than transformers with linear rather than quadratic scaling in sequence length, and the RWKV and Jamba families extend the same subquadratic direction.[^178] Sutskever's statement at NeurIPS 2024 that pretraining as we know it will end came from inside the paradigm.[^148]

The objection is that scaling remains the necessary substrate. Dibia's 2026 argument is that efficiency gains ride on capabilities first unlocked by large runs: distillation requires a large teacher, and reinforcement learning from verifiable rewards requires a base model competent enough to produce occasionally correct attempts.[^179] Epoch's own projections expect frontier training compute growth to plateau rather than reverse, and the measured 4 to 5 times annual growth continued through 2026.[^142] On that reading the efficiency turn is a change in what compute is spent on rather than a reduction in how much is spent.

Sardana and Frankle identified the specific way the optimum moves once deployment is included. Chinchilla optimality assumes training cost is the only cost. Once inference demand enters the objective, a lab serving many queries should train smaller models on far more data than Chinchilla prescribes, accepting a worse training compute tradeoff in exchange for a permanently cheaper model to serve.[^180] The paradigm's own economics therefore push toward smaller models as deployment scales, without any external constraint being invoked.

Putting the pieces together gives a specific and falsifiable expectation for training compute demand. The efficiency turn changes the composition of compute spending rather than eliminating it: pretraining on undifferentiated web text loses share to curated corpora, synthetic task generation, and reinforcement learning against verifiers, which are compute intensive in a different pattern, with heavy rollout generation and verification rather than uniform passes over a fixed corpus. The aggregate compute bill does not obviously fall. What changes is the shape of the workload and the hardware that suits it, which is where this argument connects to the infrastructure question (evidence grade: B).

Three observable events would confirm the transitional reading during the second half of 2026 and 2027: a frontier release whose headline capability gain comes with flat or reduced pretraining compute; a widening of the open to closed capability gap reversing back toward three months or less on Epoch's index; and enterprise inference volume migrating toward models under 100 billion parameters at fixed task quality. Three would refute it: continued 4 to 5 times annual frontier compute growth with proportional capability gains through 2027; the open lag widening past six months on public benchmarks; and a fixed capability price floor forming as the tenfold annual decline decelerates below a factor of three.

### 9.10 The consequence that carries into the infrastructure argument

The scaling laws are correct as empirical descriptions and they are being used as a forecast of what will remain expensive. Those are different claims.

Kaplan and Chinchilla establish that capability scales predictably with compute and data (evidence grade: A). Epoch establishes that the compute required for a fixed capability halves roughly every eight months (evidence grade: A). DeepSeek establishes that a talent dense team can reach near frontier capability at a small fraction of assumed cost (evidence grade: A for the result, B for its generality). Epoch establishes that the best open weight model trails the frontier by about four months (evidence grade: A for the measurement, B for its persistence).

Hold those four together and the position that follows is narrow. Demand for intelligence is compounding, the price of a unit of intelligence falls by roughly a factor of ten per year, and the capability that commands a premium has a depreciation half life measured in months. Capital committed to producing capability under the specific technical configuration of 2026, on a horizon long enough for two or three of those halvings to occur before the asset earns, is exposed to a change in the production function rather than to a change in demand. Section 10 examines what that exposure looks like in hardware, and Section 11 examines which project vintages carry it.

## 10. The infrastructure trap: what the buildout assumes and what the workload is doing

The preceding sections established that intelligence demand is compounding and that the wrapper layer purges itself continuously. This section examines the layer where the paper locates the genuine bubble: centralized compute infrastructure. The argument proceeds in five steps. First, "AI compute demand" is two different variables, training and inference, with different economics and different trajectories. Second, the hardware serving those variables is in the middle of a classic migration from flexible general purpose silicon to specialized silicon. Third, a parallel migration is moving a slice of inference out of data centers entirely, onto devices. Fourth, model saturation dynamics will decide how much of future token demand requires frontier hardware at all. Fifth, putting these together yields a map of where tokens will actually be generated, and that map diverges sharply from the map implied by the current construction pipeline.

### Two demands wearing one label

Epoch AI's compute allocation work formalizes what the market still prices as a single quantity: a lab that can trade compute flexibly between training and inference should treat them as two distinct optimization problems, spending comparable resources on each at a fixed performance target, with a menu of techniques (overtraining, distillation, test time compute) that convert one into the other.[^181] The two demands behave differently in every dimension that matters to an infrastructure investor. Training is lumpy, concentrated in a handful of frontier labs, and economically resembles research and development expenditure. Inference is continuous, demand driven, and linked directly to revenue. Training requires high precision arithmetic and dense interconnect between tens of thousands of chips; inference tolerates low precision and rewards specialized silicon. A **graphics processing unit** (GPU, the flexible parallel processor that has carried both workloads since 2012) can do either; the question is whether it remains the cost efficient answer for both.

The decomposition is no longer academic. Deloitte's estimate, presented at CES 2026, put inference at about 50 percent of all AI compute in 2025 with two thirds projected for 2026.[^182] Nvidia's results for the quarter reported on May 20, 2026 showed revenue of 81.6 billion dollars, up 85 percent year over year, with data center revenue of 75.2 billion dollars, up 92 percent, and management describing reasoning workloads that consume hundreds to thousands of times more tokens per task than one shot inference.[^183] Jensen Huang's accompanying release stated that "AI inference token generation has surged tenfold in just one year."[^184] Mid 2026 analyses describe the hyperscalers' 2026 capital expenditure (capex), tallied between 630 and 725 billion dollars depending on scope, as majority aimed at inference silicon, with the inference optimized chip market alone put at 50 billion dollars for 2026, and inference at 60 to 70 percent of the roughly 400 billion dollar AI accelerator market.[^185] Estimates of the inference share vary by methodology, from 50 percent to 90 percent, and Epoch AI's own accounting keeps training and experiments at a substantial share of power demand, so the split remains contested even as the direction does not.[^186]

For capital allocation the decomposition matters because the two demands can move in opposite directions. Token consumption is a measure of inference. The number of frontier training runs is a measure of something else entirely: the research strategies of fewer than ten organizations. An investor who underwrites a data center on "AI demand" has actually made two separate bets, one on the token economy and one on the persistence of brute force frontier training, and the second bet is the fragile one. Section 9 documented the efficiency forces working against it: compute for fixed performance halves roughly every 8 months,[^187] and the meek models literature argues that smaller budget models converge toward frontier performance over time, so that "even companies that can scale their models exponentially faster than other organizations will eventually have little advantage in capabilities."[^188] Through mid 2026 the frontier has kept spending anyway, with training compute still growing 4x to 5x per year,[^189] which is why this paper treats flattening training demand as a graded risk scenario (evidence grade: C) rather than an observed trend break. The point for this section is narrower: the demand that fills data centers in the 2030s is overwhelmingly likely to be serving demand, and serving demand is contestable by other silicon in a way that frontier training demand currently is not.

### The silicon cycle: flexibility first, efficiency after

Computer architecture has a standard script for what happens when a workload stabilizes. General purpose hardware wins while the algorithms churn, because flexibility is worth more than efficiency when nobody knows what next year's workload looks like. Once the workload settles, an **application specific integrated circuit** (ASIC, a chip designed for one class of computation) captures it with an order of magnitude better economics. The founding evidence for AI is Google's 2017 Tensor Processing Unit (TPU) paper, a peer reviewed study of a domain specific chip built around a 65,536 unit array of 8 bit multiply accumulate circuits delivering 92 peak tera operations per second (TOPS, trillions of arithmetic operations per second). Measured in production data centers, the TPU ran about 15 to 30 times faster than its contemporary GPU and CPU on stabilized inference workloads, with performance per watt 30 to 80 times higher, and the authors concluded that "major improvements in cost energy performance must now come from domain specific hardware."[^190] The follow up held across generations: TPU v4, published in 2023, outperformed v3 by 2.1x and improved performance per watt by 2.7x while scaling to 4,096 chips, a decade long ASIC roadmap sustained because the underlying workload, dense neural network arithmetic, had stopped changing shape.[^191]

The strongest objection to reading this as destiny comes from Sara Hooker's Hardware Lottery: specialization is a bet on workload stability, research ideas win partly because they fit incumbent hardware, and a shift in model architecture resets the ASIC clock while GPUs keep their option value.[^192] Section 12 gives that objection its full weight. What the objection cannot explain is the revealed preference of every organization with the scale to act on the question. By mid 2026, each hyperscaler runs a custom silicon program: Google's TPU v7 Ironwood reached general availability in April 2026,[^193] Meta's MTIA line, Amazon's Trainium 3, and Microsoft's Maia 200 are in deployment or ramp, and OpenAI has a custom inference chip designed with Broadcom, whose reported AI backlog stood near 73 billion dollars (third party figures, reported with caution).[^194] Industry analyses aggregated in mid 2026 put the custom AI ASIC market on a compound annual growth rate near 44.6 percent against about 16 percent for merchant GPUs, and project ASICs taking a majority of inference spend by 2027.[^194]

The economics driving the migration are visible at the token price level, with the caveat that most 2026 magnitude claims are vendor benchmarks pending independent MLPerf confirmation (MLPerf being the industry's standardized machine learning benchmark suite). Groq serves Llama 3.3 70B at roughly 750 tokens per second, priced at 0.59 and 0.79 dollars per million input and output tokens, and claims 40 to 60 percent less power than GPUs at equivalent throughput; Cerebras serves the same model at about 2,100 tokens per second at 0.85 and 1.20 dollars.[^195] Midjourney is reported to have cut inference costs 65 percent by migrating to TPUs, in a market where inference budgets are described as running 15 to 20 times training expenses for deployed applications.[^196] Ironwood itself had no public chip hour price at general availability, and comparison pieces lean on vendor throughput claims, so the honest statement is that the peer reviewed record establishes the order of magnitude ASIC advantage on stable workloads, while the 2026 magnitude estimates carry a B grade pending independent audits.[^193]

The implication for the data center layer is direct. A facility whose financial model assumes GPU class power density, GPU class rental pricing, and GPU class residual values is exposed to a serving market that migrates to silicon with several times better performance per watt. The migration does not reduce the number of tokens generated. It reduces the number of GPU hours, and the megawatts, needed per token, and it moves the margin to whoever owns the specialized silicon rather than to whoever owns the generic building around it.

### The device layer: an installed base forming at consumer electronics speed

The second migration route bypasses the data center entirely. A **neural processing unit** (NPU, a low power inference accelerator integrated into consumer devices) now ships in a large and rising share of phones and laptops. Counterpoint Research forecasts generative AI capable smartphones at 45 percent of global shipments in 2026, up from 36 percent in 2025, heading to 52 percent in 2027, led by Apple and Samsung premium portfolios.[^197] The forecast is more striking against the backdrop of a shrinking overall market: global smartphone shipments are projected to decline 13.9 percent year over year to 1.08 billion units in 2026 while the AI capable share rises.[^198] On the PC side, Copilot Plus class devices crossed 42 percent of new Windows notebook shipments by June 2026 as Snapdragon X2 ramped.[^199] Apple had shipped over 450 million Apple Intelligence capable devices through the first quarter of 2026.[^200] The academic literature has caught up with the hardware: systematic reviews of on device language models document meaningful workloads running on phone class NPUs through compression and quantization,[^201] and the mobile edge inference literature positions local serving as more cost effective, latency efficient, and privacy preserving than the cloud for a wide class of tasks.[^202]

The bounds on this migration are physical. Shipping NPU TOPS figures overstate usable throughput, and memory bandwidth rather than arithmetic capacity caps on device model quality, which keeps frontier scale inference in data centers.[^203] Batch serving economics also favor the center, a point Section 12 develops. The honest summary of the evidence is that device migration is currently a capacity story rather than a measured traffic story: the installed base is verifiably enormous and growing, while no source yet measures how much cloud inference volume it has displaced (evidence grade: B).[^201] The emerging architecture is hybrid routing, in which each request is placed on the cheapest adequate tier, device, edge, or cloud, by complexity, latency, and cost, described in 2026 practice guides as the default pattern.[^204] Hybrid routing does not empty data centers. It skims the commodity slice, transcription, summarization, autocomplete, photo processing, exactly the high volume low complexity traffic that utilization models for generic inference capacity implicitly count on.

### Saturation: the smartphone template

The third force in the trap is behavioral. Clayton Christensen's performance oversupply framework predicts that when product performance exceeds what the mainstream can absorb, the basis of competition shifts from performance to reliability, then to convenience, then to price.[^205] Herbert Simon's satisficing result supplies the micro foundation: agents stop optimizing once an aspiration level is met.[^206] The consumer electronics record shows what this does to upgrade economics. The global smartphone replacement cycle stretched from about 2.4 years to about 3.5 years as devices became good enough, with 75 percent of upgraders citing battery and 55 percent citing screen damage rather than new features as the trigger,[^207] and reporting from August 2026 expects the cycle to stretch to four years, driven by rising component prices and incremental hardware gains.[^208] Forty two percent of US iPhone buyers in 2026 were replacing a phone owned three years or longer, up from 32 percent a year earlier.[^209]

The AI market is already tracing the same curve at the token level. On OpenRouter, the largest neutral routing marketplace, the five fastest growing models in April 2026 all offered free access or pricing below 1 dollar per million tokens, Chinese origin models exceeded 45 percent of the roughly 20 trillion tokens routed weekly, and the provider share table put Xiaomi at 22.3 percent of routed tokens against OpenAI at 8.1 percent: frontier brand share is not usage share.[^210] This is the pattern that versioning economics predicts. Deneckere and McAfee's damaged goods analysis established that sellers profitably serve the mass market with cheaper degraded tiers while premium buyers self select upward, and that the two tier outcome can leave everyone better off,[^211] which is precisely the frontier versus mini and legacy tiering the labs adopted in 2025 and 2026. Meanwhile Publicis Sapient's 2026 enterprise survey found 47 percent of enterprises judging AI already fully capable of meeting today's business needs, with 42 percent saying their organizations are not built to capture the value,[^212] a direct measurement of a good enough segment forming at enterprise scale.

Who keeps paying for the frontier? Eric von Hippel's lead user theory predicts a small segment whose needs run ahead of the mainstream by months or years,[^213] and the revenue data matches: Anthropic's usage concentrates in developer and technical workloads, with coding around 36 percent of Claude.ai usage in late 2025 and roughly 80 percent of revenue from enterprise API consumption per third party trackers reporting about 30 billion dollars of annualized recurring revenue (ARR) by April 2026, figures not officially confirmed.[^214] The counterexample that keeps the claim honest is OpenAI, where roughly 70 percent of revenue comes from consumer subscriptions,[^215] proof that a frontier business can be mass funded. The saturation claim is therefore graded, and bounded: the frontier tier retains a paying market, while the mass of token volume migrates to tiers that run happily on older, cheaper, more specialized hardware. Two open questions cut against saturation, and Section 12 takes both seriously: production agents still fail roughly one in three structured benchmark attempts, so the reliability bar has not been cleared even where the capability bar has,[^216] and interface improvements could re expose mainstream users to frontier capability, re triggering upgrades the way cameras once did for phones.[^217]

### Where tokens will be generated

Assemble the four pieces and the shape of the trap appears. Token demand is exploding: Google processed 3.2 quadrillion tokens per month by May 2026, against 480 trillion a year earlier and 9.7 trillion in 2024, a roughly 330x rise in two years,[^218] while OpenRouter routed 28.9 trillion tokens in the peak week of May 18 to 24, 2026, up 7.4 percent week over week.[^219] OpenAI served roughly 900 million weekly active users with its chief financial officer telling employees that July 2026 ARR topped the entire second quarter.[^220] The boom in intelligence consumption is beyond dispute. The contested question is what fraction of those tokens, five years out, is generated on merchant GPUs inside newly built centralized data centers, because that fraction, and only that fraction, services the debt on the current construction pipeline.

The forces pushing the fraction down are the three documented above: ASIC serving with multiples better performance per watt, an on device installed base approaching half of all new consumer hardware, and a demand distribution already majority weighted toward cheap tiers that need none of the frontier's infrastructure. The forces holding the fraction up are real and are treated fully in Section 12: reasoning workloads that multiply tokens per task by orders of magnitude, batch economics that no single device can match, and model sizes above device memory. The academic case that this contest could break centralized data center economics exists in the mobile edge inference literature, and this paper grades it honestly: the mechanism is plausible and its ingredients are verified, while no source yet demonstrates realized centralized demand destruction, so the strong version of the claim carries an analogical grade (evidence grade: C).[^221] Epoch AI's forecast that inference compute passes training compute by 2030, with a large share moving to ASICs,[^222] and Grogan's 2026 argument that the economic center of AI is shifting from pretraining to inference as infrastructure,[^223] both point the same direction: the buildout is financing yesterday's workload map.

What the trap does not require is any decline in AI usage. That is the crux of the whole paper's thesis, and the reason the coming years will be misread. A facility financed in 2025 on the assumption that tokens equal GPUs equal centralized megawatts can fail financially while every input to that equation grows, because each link in the chain, token to GPU, GPU to data center, data center to new construction, is being weakened independently by specialization, saturation, and distribution. The next section prices that failure mode and dates it.

## 11. Stranded assets and timing: pricing the mismatch

A **stranded asset**, in Ben Caldecott's canonical definition, is an asset that has suffered unanticipated or premature write-downs, devaluations, or conversion to liabilities.[^224] The concept was developed for fossil infrastructure, and the follow up literature stresses that stranding occurs regularly as an ordinary feature of economic development, well beyond climate policy.[^225] Applied to compute, the definition translates cleanly: a data center strands when the financial assumptions embedded at signing (hardware value, rental rates, utilization, power density) move faster than the physical asset can adapt. This section measures the speed differential, examines the accounting battle already being fought over it, sizes the pipeline exposed to it, maps the financing structures that amplify it, documents the first realized casualty, and makes the vintage call the brief demands.

### The clock mismatch

Three clocks run at different speeds. The construction clock: hyperscale build cycles run roughly 2 to 4 years from commitment to energization.[^226] The grid clock: interconnection waits in Northern Virginia, Phoenix, and Dallas now run 4 to 7 years, captured in the industry aphorism that in Northern Virginia "the grid wait is seven years, while the building takes two,"[^227] with substation transformer lead times exceeding 160 weeks in 2026, up from about 150 in 2025 and 140 in 2023.[^228] And the silicon clock: Nvidia has moved to an annual architecture cadence, reconfirmed across CES 2026 and GTC 2026, with Blackwell Ultra in the second half of 2025, Vera Rubin in the second half of 2026, Rubin Ultra in the second half of 2027, and Feynman in 2028.[^229]

A project signed in 2025 for delivery in 2028 or 2029 therefore commits to power density, cooling architecture, and rack layout across two to four silicon generations it has not seen. Each Nvidia generation has raised per rack power draw, and a shell engineered around one density assumption can be physically unable to host the racks of the generation that ships when it opens. The Harvard Belfer Center's February 2026 grid study documents the electrical side of the same mismatch: the grid investments AI load requires operate on decade scale planning cycles against a compute industry replanning annually.[^230] Analysts who track the cadence itself have started asking whether the annual rhythm is an innovation engine or, in one January 2026 formulation, capex quicksand for the customers locked into it.[^229]

### Depreciation: the Burry critique and the waterfall defense

The accounting expression of the clock mismatch is the depreciation schedule, and the fight over it is the most important accounting dispute in the market today. Michael Burry's estimate is that hyperscalers will understate depreciation by roughly 176 billion dollars over 2026 to 2028, overstating profits by more than 20 percent, by depreciating accelerators over 6 years while the frontier obsolescence cycle runs 2 to 3.[^231] The schedule history shows how much room managements have found in this judgment. Microsoft extended servers from 4 to 6 years in 2022, Alphabet moved to 6 years in January 2023, and Meta extended to 5.5 years effective January 1, 2025. Amazon then broke the pattern in February 2025, shortening a subset of servers and networking equipment to 5 years and citing the pace of AI innovation; the effect for the nine months ended September 30, 2025 was 889 million dollars of additional depreciation and 677 million dollars less net income.[^232] One company moving one year on one subset of assets produced a nine figure earnings swing, which calibrates what an industry wide correction toward Burry's assumption would do.

The defense deserves its full weight. State Street's argument is that the critique conflates two different cycles: the frontier training obsolescence cycle of 2 to 3 years, and the economic utility cycle of the hardware waterfall, in which accelerators cascade from frontier training to inference to internal workloads, making 4 to 6 year useful lives defensible.[^233] Nvidia makes the same argument, reporting that customers observe 4 to 6 year depreciable lives based on utilization and longevity.[^234] The waterfall is real, and it is precisely why this paper's stranding argument runs through the inference layer rather than around it: the waterfall's terminal stages assume that cascaded GPUs remain the cost efficient way to serve inference. Section 10 documented the ASIC and edge forces attacking exactly that assumption. If cascaded H100s must compete for serving workloads against ASICs with multiples better performance per watt and against NPUs that serve the commodity slice at the device's marginal cost, the waterfall's later years yield less revenue than the schedules assume, and Burry's arithmetic re enters through the back door even where his premise about training obsolescence is successfully rebutted. A functioning secondary market for used accelerators already exists and prices this continuously,[^235] which makes the residual value assumption testable quarter by quarter rather than a matter of managerial assertion.

### 190 gigawatts announced, 5 under construction

The scale of the exposed pipeline is documented by the buildout's own tracking. Since 2024, developers have announced approximately 190 **gigawatts** (GW, billions of watts of electrical capacity) of future data center capacity; of the 16 GW slated for 2026, only about 5 GW is actually under construction.[^236] Between 30 and 50 percent of 2026 US data center capacity is expected to be delayed or canceled, and more than 150 billion dollars of US projects were blocked or delayed in 2025.[^237] Capex of the largest data center firms nears 750 billion dollars in 2026, up from roughly 450 billion in 2025, and analyst expectations for the 14 largest developers' 2027 spending climbed 56 percent between August 2025 and February 2026.[^236] As of May 2026 the gating factor is no longer capital or chips; it is the physical electrical layer, high voltage transformers and medium voltage switchgear.[^236]

Stargate, the flagship of the announcement era, illustrates the gap between the announced and the actual. By late 2025 it had reached about 7 GW of planned capacity with over 400 billion dollars committed; in March 2026, parts of the OpenAI and Oracle expansion were scaled back amid financing complexity and revised demand assumptions. As of mid 2026 the project has 0.3 GW operational at Abilene, six additional US sites under construction, and over 9 GW projected by 2029, with the Michigan 1 GW campus on phased delivery through 2027.[^238] None of this is failure; 0.3 GW of operating frontier capacity is a real asset. It is, however, a 30 to 1 ratio between announcement and operation at the single most publicized project in the industry, and announcement to completion gaps of 2 to 4 years mean that capacity announced at peak enthusiasm delivers into whatever hardware paradigm and pricing environment exists at completion.

The historical template for late cycle infrastructure commitments is Odlyzko's study of the British railway mania, the largest technology investment mania on record relative to national income, whose collapse hit hardest those who subscribed capital at the top.[^239] Campbell's work adds the financing detail that matters most for the analogy: railway shares fell about 50 percent between 1846 and 1850, and losses were compounded when companies called the remaining 90 percent due on partly paid shares, so late subscribers carried open ended capital commitments into the bust.[^240] The modern equivalent of the uncalled share is the multi year take or pay lease and the phased capex commitment: obligations that continue drawing capital after the thesis that justified them has broken. Odlyzko's companion study of the 1860s mania shows the pattern repeating one cycle later through new financing instruments,[^241] which is the correct prior for evaluating today's structures.

### Financing structures: the leverage map

Three structures concentrate the risk. First, the neocloud model: specialized GPU rental companies financed with debt collateralized by the GPUs themselves and by contracted revenue from a small number of counterparties. The collateral depreciates on the silicon clock documented above, so the loan to value ratio deteriorates with every Nvidia keynote, and the secondary market that would absorb liquidated collateral prices used accelerators continuously against newer generations.[^235] Second, vendor and circular financing: the Bank for International Settlements' systemic risk study, reported in July 2026, found the AI boom larger than every historical technology bubble on its metrics and identified circular financing arrangements, hyperscaler equity stakes funding labs that buy compute back from the same hyperscalers, as a closed valuation loop that couples the layers, so a shock in one propagates to all.[^242] Third, special purpose vehicles that isolate single tenant facilities: efficient in calm markets, they concentrate counterparty risk exactly where tenant default or renegotiation is most plausible, in facilities whose alternative use value is lowest. The most exposed physical assets, per mid 2026 analysis, are generic powered shells in weak locations, cheap land with constrained transmission and poor fiber, whose only bid came from the boom itself.[^243]

### The first realized casualty: Situational Awareness LP, July 2026

The thesis of this paper predicts that financial vehicles crack before the technology does. In late July 2026, the prediction acquired its first major realized instance. Situational Awareness LP, the AI concentrated hedge fund founded in 2024 by Leopold Aschenbrenner, former OpenAI researcher and author of the June 2024 essay of the same name, grew from a few hundred million dollars at launch past 20 billion dollars of assets under management, with a roughly 270 percent net gain in 2026 through May and more than 1,000 percent since inception;[^244] the Financial Times reported a 439 percent net return for the first half of 2026, and assets peaked near 45 billion dollars in early July.[^245] The book was concentrated leveraged longs in AI infrastructure, SK Hynix, CoreWeave, Nebius, Micron, Bloom Energy, with reported leverage up to 400 percent.[^246]

Then the infrastructure complex repriced. The Philadelphia Semiconductor Index fell 28.6 percent from its June 22, 2026 peak; the Morgan Stanley Momentum TMT Index dropped 53.5 percent in July; the fund's concentrated holdings fell between 35 and 47 percent in the month.[^247] Margin calls from prime brokers Goldman Sachs, JPMorgan, and Bank of America forced the fund to sell its entire public equity portfolio to Citadel at a discount.[^248] Assets collapsed from about 45 billion to about 10 billion dollars, with the remainder dominated by a roughly 5 billion dollar private stake in Anthropic, and the fund reported as continuing as a private vehicle.[^249] SpotGamma's post mortem describes a four sided squeeze: longs down 35 to 47 percent, shorts moving against the book, margin calls, and no fresh capital.[^250] Some circulating figures, including a claimed 67 percent drawdown and the exact Citadel discount, were not independently verified at publication, and this paper does not rely on them; the forced unwind, the Citadel purchase, and the collapse in assets are multi sourced across CNBC, Bloomberg, and the FT.[^251]

The instructive feature is that the fund's thesis had been directionally right, spectacularly so, through the first half of the year. It failed on leverage and crowding, a liquidity failure inside a directional win. That is the capital cycle signature this paper's historical sections predict: the first casualties of an infrastructure repricing are the leveraged vehicles built on the infrastructure trade, while the underlying usage curve, 3.2 quadrillion tokens a month and rising through the exact weeks of the unwind, does not flinch.[^218] One fund is one fund; the reading offered here, that July 2026 marks the turning point of the cycle rather than an idiosyncratic blowup, is a hypothesis the observable sequence in Section 13 will confirm or refute.

### The vintage call

The brief requires a call, so here it is, with its falsification conditions. Safe vintages: capacity operating or energizing in 2026 and 2027 that is pre leased to investment grade tenants, sited on secured power with interconnection already granted, and engineered with density and cooling headroom (liquid cooling ready, high ceiling power provisioning). This capacity delivers into a market where, as of mid 2026, vacancy is near record lows, demand exceeds deliverable supply, and the binding constraint is transformers rather than tenants.[^252] Its risk is margin compression late in life, in the 2030s serving market. Exposed vintages: projects announced from mid 2025 onward for delivery in 2028 and 2029, financed on merchant or single tenant assumptions, in weak locations, with GPU collateralized debt or take or pay obligations extending past 2030. These deliver two to four silicon generations late into a serving market being restructured by ASICs and edge inference, and they carry the railway mania financing signature: capital calls that outlive the thesis.[^240] Most exposed of all: generic powered shells without secured interconnection, whose entire value case was extrapolated scarcity.[^243] Falsifiers for this call: if frontier training compute keeps compounding at 4x to 5x per year through 2028,[^189] if reasoning token demand keeps outrunning ASIC serving capacity, and if pre leasing at investment grade covenants extends to the 2028 and 2029 cohort, then the late vintages get bailed out by demand and the call fails. Each of those is observable quarterly.

### The fiber template, with its honest disanalogy

The closest historical template is the telecom fiber overbuild of 1998 to 2001. The financing myth was quantified: the industry promoted internet traffic doubling every 100 days, roughly 1,000 percent annual growth, while Odlyzko and Coffman measured actual growth at 75 to 150 percent per year.[^253] Roughly 2 trillion dollars built 80 to 90 million miles of fiber, 95 percent of it dark by 2001;[^254] only an estimated 2.7 to 5 percent of installed capacity was lit even years after the bust, bandwidth prices collapsed by up to 90 percent, and the sector produced more than 60 bankruptcies, WorldCom and Global Crossing among them.[^255] The Federal Reserve's post mortem treats it as a capacity boom and bust in the classic mold.[^256] Meanwhile traffic kept doubling roughly annually straight through the crash: usage and investment returns fully decoupled,[^257] and the dark fiber was eventually lit, becoming the physical substrate of broadband, streaming, and cloud, a socially productive overbuild whose private investors were nonetheless destroyed.[^258]

Today's demand myth candidate is the straight line token extrapolation, and today's counter evidence is stronger than 2000's: capacity is pre leased years ahead by investment grade tenants and the mid 2026 bottleneck is delivery rather than demand,[^252] a configuration absent from the 2000 setup. The honest disanalogy cuts the other way. Fiber glass had a 20 to 30 year physical life, so it could wait dark for demand to arrive; GPU silicon obsoletes on a 1 to 3 year architecture cadence, so an AI overbuild strands faster and reuses worse at the chip layer.[^259] What carries over to the deployment phase this time is the durable perimeter, shells, power infrastructure, cooling plant, grid interconnection rights, while the accelerators inside are consumables. The fiber template therefore predicts the shape of the bust, financing vehicles dying while usage compounds, and simultaneously warns that the post crash bargain hunting will be in sites and megawatts rather than in the chips (evidence grade: C for the template as a whole).[^260]

## 12. Counterarguments, steelmanned then answered

A thesis that cannot survive its strongest objections is a mood. This section takes the nine most serious objections to the paper's position, states each in its best form with its best evidence, and then answers it. In order: the Hardware Lottery and architecture churn; batching and test time compute favoring centralization; device memory ceilings; Jevons absorbing every efficiency gain; frontier training runs still growing; the revenue gap indicting the whole stack; circular financing fragility; vendor benchmark reliance; and the AlphaChip reproduction dispute.

### Objection 1: architectures change too fast for ASICs

The steelman. Sara Hooker's Hardware Lottery argument is the deepest version: research ideas succeed partly because they fit incumbent hardware, specialization is a bet on workload stability, and domain specific silicon carries obsolescence risk whenever model architectures shift.[^192] The AI research frontier churns visibly: subquadratic architectures (Mamba, RWKV, Jamba), new attention variants, new modalities, and Sutskever's declaration at NeurIPS 2024 that "pretraining as we know it will end" all signal that the workload has not stabilized.[^261] An ASIC taped out for 2026's transformer could be a paperweight against 2028's architecture, and the GPU's option value is exactly what a period of churn rewards. On this view the ASIC migration thesis repeats the mistake it accuses the GPU buildout of making: pricing in the permanence of a paradigm.

The answer concedes the mechanism and disputes its scope. The churn that matters for serving silicon is churn in the deployed workload, and the deployed workload has stabilized around transformer class inference to a degree the research frontier has not: a decade of TPU generations, v1 through the v7 Ironwood that reached general availability in April 2026, sustained order of magnitude efficiency gains precisely because dense neural inference stopped changing shape underneath them.[^191] Every hyperscaler has examined Hooker's risk with its own capital and concluded the workload is stable enough to tape out against; Google, Amazon, Meta, Microsoft, and OpenAI with Broadcom are collectively spending tens of billions on that judgment.[^194] The risk allocation matters more than the risk itself: hyperscalers deploying ASICs into their own serving fleets internalize the obsolescence risk and can absorb a reset, while the paper's stranding argument concerns third parties who financed generic GPU capacity on the assumption that no specialization would occur at all. And if Hooker is right and a post transformer architecture resets the ASIC clock, the reset strands the current GPU fleet with equal force, so the objection weakens the buildout's economics from another direction rather than rescuing them.

### Objection 2: batching and test time compute keep inference centralized

The steelman, which is the strongest objection in this list. Batch serving economics structurally favor the center: amortizing model weights and hardware across thousands of concurrent requests is a cost lever no single user device possesses, and frontier models exceed device memory outright.[^262] Test time compute makes the pull stronger, because reasoning models convert inference tokens into accuracy, consuming hundreds to thousands of times more tokens per task than one shot inference, per Nvidia's May 2026 commentary, with inference token generation up tenfold in one year.[^184] Reasoning workloads are the fastest growing inference segment and cannot run on a phone.[^263] On this evidence, the token future is a heavy infrastructure future: large, power hungry, centrally operated fleets, exactly what the buildout is building.

The answer accepts most of the premise. Centralized inference volume will grow enormously; this paper's thesis requires no other outcome, since its demand side argument (Sections 6 through 8) predicts explosive token growth. The objection fails as a defense of the exposed assets because it defends the wrong variable. The stranding argument concerns which silicon, owned by whom, in which facilities, captures the centralized volume. Reasoning tokens served on TPU v7 pods, Trainium fleets, or Groq and Cerebras hardware are centralized tokens that generate zero revenue for a merchant GPU neocloud in a weak location; the batch economics the objection invokes are available to every ASIC operator, and at 60 to 70 percent of accelerator spending already going to inference,[^185] the specialized silicon is arriving exactly where the batches are. Centralization of tokens and stranding of late vintage GPU projects are compatible outcomes, and the mid 2026 data shows both trends running concurrently: tenfold reasoning token growth and a custom ASIC market compounding near 44.6 percent annually against 16 percent for merchant GPUs.[^194] The objection also proves less than it seems over time: test time compute is a technique for buying accuracy with tokens, and the efficiency literature's 8 month halving of compute per fixed capability applies to reasoning chains as it applied to pretraining.[^187]

### Objection 3: device memory ceilings cap the edge migration

The steelman. Marketing TOPS overstate usable on device throughput; memory bandwidth, and memory capacity, cap the size and quality of models a phone or laptop can run, and the gap between a device class model and a frontier data center model is not closing from the device side.[^203] The 45 percent GenAI smartphone shipment share is a hardware attach rate, an accounting of chips shipped rather than of workloads moved;[^197] no published source measures actual cloud inference displacement, and hybrid architectures like Apple's Private Cloud Compute explicitly route hard queries back to the center. The edge migration could remain what it is in mid 2026: an impressive installed base serving autocomplete.

The answer is to accept the bound and price what remains inside it. The paper's Section 10 already grades the edge claim at B and calls it a capacity story rather than a traffic story; nothing in the stranding argument requires devices to serve frontier workloads. What devices capture is the commodity slice, transcription, summarization, photo processing, short form generation, and the saturation data shows the commodity slice is where the token mass lives: the fastest growing routed models are free or under a dollar per million tokens.[^210] Utilization models for generic inference capacity are calibrated on serving everything; losing the high volume floor to silicon the user already paid for lowers realized utilization and pricing power at the margin, which is where debt service lives. The objection also ages badly on its own terms. Device memory is the binding constraint today, and the on device literature documents compression, quantization, and distillation moving capability into fixed memory budgets year over year,[^201] while the models that count as good enough keep getting smaller (Section 9's efficiency evidence). A ceiling that rises annually against a sufficiency bar that falls annually is a narrowing moat for centralized commodity serving.

### Objection 4: Jevons absorbs every efficiency gain

The steelman. The paradox named for Jevons, formalized in Alcott's treatment, holds that efficiency gains lower effective cost and raise total consumption; Jevons's own line is that "it is a confusion of ideas to suppose that the economical use of fuel is equivalent to diminished consumption. The very contrary is the truth."[^264] Satya Nadella invoked exactly this for AI in January 2025.[^265] The measured elasticity supports it: a 1 percent decrease in compute price generates a 1.42 percent increase in volume, and Goldman Sachs projects global token consumption multiplying 24x between 2026 and 2030.[^266] On this view every ASIC, every distillation, every NPU makes intelligence cheaper, demand more than compensates, and the data centers fill regardless. Efficiency is the bull case for infrastructure.

The answer is that Jevons governs the quantity of consumption while remaining silent on where the consumption lands, and the second variable is the one infrastructure returns depend on. Alcott's formalization contains the answer: consumption migrates to whatever form is cheapest, which need not be the incumbent infrastructure.[^264] The clean historical test is telecom. Internet traffic obeyed Jevons perfectly, doubling roughly every year through the 2001 to 2003 bust without interruption, and the owners of the centralized capacity built for that traffic went bankrupt in their dozens, because the traffic growth arrived at collapsed prices on consolidated infrastructure.[^257] The 2026 configuration repeats the setup: tokens compounding at 330x over two years at Google[^218] while serving migrates to ASIC pods and edge silicon that generate those tokens at a fraction of the GPU cost per token. A super elastic demand curve at collapsing unit prices is a wonderful world for consumers of intelligence and a treacherous one for owners of the previous generation's capacity, since revenue per facility is price times captured volume, and both terms move against the late vintage GPU project. Jevons rescues the aggregate; the aggregate was in no danger. The claim that token consumption can grow 100x to 1000x while demand for new centralized GPU data centers simultaneously collapses is exactly the conjunction the telecom precedent realized, and it carries a B grade here: both halves are separately documented in 2026, the conjunction is not yet observed, and mid 2026 centralized demand still exceeds deliverable supply.[^252]

### Objection 5: frontier training keeps growing, so training demand rescues the buildout

The steelman. The realized data through mid 2026 shows no flattening: Epoch AI measures frontier training compute growing 4x to 5x per year,[^189] the second quarter of 2026 saw simultaneous Grok 5, GPT-5.5 class, and DeepSeek V4 cycles described as the most compute and money spent in any quarter of AI history,[^267] single frontier runs cost 200 to 500 million dollars at 10^26 to 10^27 floating point operations with 1 to 3 billion dollar runs projected for late 2027,[^268] and Epoch expects a plateau rather than a downturn even in conservative scenarios.[^189] Efficiency has so far behaved as a Jevons effect on training itself: cheaper capability raised total training spend. The meek models thesis is a projection; the spending is a fact.

The answer concedes the fact and examines its structure. This paper grades flattening training demand at C, an analytically grounded scenario rather than an observed break, and asserts nothing stronger. What the objection cannot deliver is a rescue of the exposed pipeline, for three structural reasons. First, concentration: frontier training demand comes from fewer than ten organizations, each vertically integrating toward owned or hyperscaler partnered capacity (Stargate for OpenAI, TPU fleets for Google DeepMind, Trainium for Anthropic's training mix), leaving merchant capacity as the residual buyer's market. Second, lumpiness and mobility: a training run is a project, negotiated and relocatable to wherever power is cheapest, offering none of the annuity revenue that debt service requires; a facility underwritten on training demand holds a pipeline of bids, and holds it against counterparties with overwhelming negotiating leverage. Third, arithmetic: even at 4x to 5x annual frontier growth, training is the minority and shrinking share of accelerator spend, 30 to 40 percent in 2026 decompositions and falling,[^185] and the sub frontier mass of training demand is collapsing in unit cost, with fixed capability targets getting roughly 10x cheaper per year to hit.[^187] The 190 GW announcement pipeline cannot be filled by ten customers running lumpy projects; it requires the inference annuity, which returns the argument to Objections 2 and 4.

### Objection 6: the revenue gap indicts the whole stack, model layer included

The steelman. David Cahn's 2024 arithmetic identified a 600 billion dollar annual hole between AI infrastructure spend and the end user revenue needed to justify it, at a time when the largest AI company earned 3.4 billion dollars annualized;[^269] by 2026 the industry spends roughly 400 billion dollars a year on compute against 50 to 60 billion of AI revenue, and hyperscaler capex near 750 billion runs an order of magnitude ahead of the AI revenue base.[^236] Cornell and Damodaran's big market delusion supplies the theory: when a market is perceived as enormous, every participant is priced to win a large share, so the sector's aggregate capitalization exceeds any consistent view of the market, and individually defensible valuations are collectively irrational.[^270] McKinsey's survey adds that only 39 percent of adopting organizations report enterprise level earnings impact.[^271] On this view the paper's distinction between a sound model layer and a bubbly infrastructure layer is a distinction without a difference: the same missing revenue underlies both.

The answer runs through the trajectory and the incidence of the gap. The gap is real and this paper does not price it away; the question is which layer's economics it breaks. Model layer revenue is compounding at rates without precedent in enterprise software: OpenAI from roughly 5.7 billion dollars in the first quarter of 2026 toward a 30 billion dollar full year target, with July ARR exceeding the entire second quarter;[^272] Anthropic's annualized revenue rose from about 9 billion dollars at end 2025 to about 47 billion by May 2026 per third party profiles, a figure reported gross of cloud reseller payouts and flagged accordingly.[^273] Consumer willingness to accept compensation for losing AI chatbots reached a mean of 124.50 dollars per month in March 2026, above the average US household mobile bill, with the median far lower at 11.40 dollars, so the value distribution is skewed toward heavy users but growing.[^274] A revenue base doubling or better annually closes Cahn's hole on a visible trajectory; what it cannot do is close the hole on the timetable of the debt financed 2028 and 2029 delivery cohort, which is precisely the vintage distinction Section 11 draws. Cornell and Damodaran's mechanism, meanwhile, is a mispricing of equity dispersion within a genuinely large market, their examples being sectors where the technology succeeded while most tickers failed; applied to AI it predicts exactly this paper's conclusion, a real market with concentrated winners and a purge of the collectively overpriced periphery, and it says nothing that rescues the aggregate infrastructure bet. The revenue gap indicts timing and leverage. Those live in the infrastructure and financing layers.

### Objection 7: circular financing makes labs and infrastructure jointly fragile

The steelman. The Bank for International Settlements' systemic study found the AI boom exceeding every historical technology bubble on its metrics and identified the circular financing loop, hyperscalers taking equity stakes in labs that spend the proceeds buying compute from the same hyperscalers, as a closed valuation circuit: revenue, valuation, and capex validate one another, so a shock at any node propagates to all of them at once.[^242] The same structure could synchronize the wrapper purge into a single funding crash.[^242] If the layers are financially coupled, the paper's layered risk decomposition is an accounting fiction, and "data center bubble" versus "AI bubble" is a semantic preference.

The answer is that the objection identifies a real amplifier and misidentifies what it amplifies. Circularity couples balance sheets; it does not homogenize the underlying assets. When the loop deleverages, each layer falls back on its own fundamentals: the labs fall back on 900 million weekly users and compounding token revenue,[^220] the hyperscalers on cash generative core businesses, and the late vintage merchant infrastructure on utilization models that Section 10 and 11 showed to be the weakest fundamentals in the system. The July 2026 sequence ran the experiment at fund scale: the leveraged financing vehicle was destroyed, the infrastructure equities it held repriced 35 to 47 percent, and the usage curve and lab revenue did not register the event.[^247] A financing shock travels the loop and settles where the fundamentals are thinnest, which is the paper's thesis restated through the transmission mechanism. The BIS finding therefore sharpens rather than refutes the layered reading: it explains why the coming repricing will look system wide for a quarter or two, and why the recovery will sort survivors by layer. The residual truth in the objection is real and worth stating plainly: labs holding multi year compute purchase commitments contracted at peak pricing carry infrastructure risk onto their own balance sheets, and the labs most exposed to that channel deserve the discount the market will apply.

### Objection 8: the ASIC case rests on vendor benchmarks

The steelman. The 2026 efficiency claims driving the migration narrative are overwhelmingly vendor sourced: Groq's power figures, Cerebras's throughput, claimed TPU generation over generation gains, all without independent MLPerf confirmation; Ironwood reached general availability in April 2026 with no public chip hour price, so nobody outside Google can compute its true cost per token; and CUDA generation gains on B200 class parts keep closing the measured gap in the public MLPerf cycles.[^193] An investment thesis built on marketing decks deserves an evidence discount.

The answer is to accept the discount and show the thesis survives it. Strip out every vendor claim and the ASIC case retains its peer reviewed core: the 2017 TPU paper's production measurements, 15 to 30x speedups and 30 to 80x performance per watt against contemporary general purpose parts on stabilized inference,[^190] and the 2023 TPU v4 paper's sustained 2.7x per watt generational gain,[^191] both published at the field's flagship architecture venue with a decade of production deployment behind them. On top of the published record sits revealed preference, the costliest signal in the market: five hyperscalers funding custom silicon programs at multi billion dollar scale, Broadcom's design win backlog near 73 billion dollars, and customers like Midjourney migrating production workloads for a reported 65 percent cost reduction.[^194] Organizations with perfect internal visibility into their own serving costs are betting their capex on the ASIC answer. The prudent reading, adopted throughout this paper, grades the physics and the direction A on the peer reviewed record and the 2026 magnitudes B pending independent audits, and no load bearing conclusion in Sections 10 or 11 requires the vendor numbers to be exactly right; the stranding argument needs specialization to capture stable serving workloads at materially better economics, a claim the published record established in 2017.

### Objection 9: the AI designed hardware loop is disputed at its foundation

The steelman. The paper's hardware specialization story leans partly on AI accelerating chip design, and its flagship evidence is contested. Igor Markov's False Dawn reevaluation argues the 2021 Nature AlphaChip paper withheld critical methodology and inputs, and independent reproductions in the Cheng and Kahng line found the reinforcement learning placer did not clearly beat strong existing heuristics; the dispute ran through Communications of the ACM, the field's journal of record.[^275] If the acceleration claim rests on an unreproducible benchmark, the "AI designs its own successor silicon" loop is a story, and the D graded claim that model specific ASICs could iterate like consumer electronics collapses with it.

The answer separates what is disputed from what is deployed. The Markov and Cheng critiques target the magnitude of the reinforcement learning placer's advantage over academic baselines on benchmark circuits; Google's production use is undisputed, and the October 2024 Nature addendum documents the methodology shaping successive TPU generations, identifying pretraining as the ingredient the reproductions omitted.[^276] Independently of Google entirely, the two dominant electronic design automation (EDA) vendors report production adoption at scale: Synopsys DSO.ai with over 700 production tape outs, Cadence Cerebrus with over 1,000 completed designs, with named customers (Samsung, SK hynix, Renesas, MediaTek) reporting consistent power, performance, and area improvements, and agentic AI workflows announced with Nvidia at the 2026 Design Automation Conference.[^277] Seven hundred commercial tape outs do not depend on one contested Nature benchmark. The paper's own grading already absorbs the objection's force: AI accelerating production chip design carries an A grade; the stronger claim that this materially shortens the model to custom silicon loop carries a B, because mask making, tape out queues, and fab lead times still dominate calendar time; and the consumer electronics iteration scenario for ASICs is explicitly graded D and presented as a hypothesis, with leading edge yields (SMIC's DUV based 5 nanometer pilot reportedly near 20 percent) arguing the opposite for now.[^278] The thesis uses the disputed claim at the weight the evidence supports and no more.

## 13. Synthesis and positioning

The paper's argument is now complete in its parts. This final section assembles it into positions: which actors the evidence says are durably advantaged, which are threatened, which are mispriced sleepers; the observable sequence of events that would confirm or refute the thesis; and the phoenix scenario that describes the cycle's end state. Each name below is carried by a mechanism documented in the preceding sections, and every position stated here is conditional on the falsifiers listed with the vintage call in Section 11.

### Durable winners

The durable winners own one of three things: the specialized silicon that serving migrates toward, the physical scarcities that survive any silicon transition, or the demand relationship that survives any infrastructure repricing.

On silicon, Google holds the strongest integrated position: a decade of TPU generations culminating in Ironwood's April 2026 general availability, vertical integration from AlphaChip assisted design through owned serving fleets, and the largest measured token flow in the world at 3.2 quadrillion per month.[^218] Broadcom wins as the design partner of the custom silicon era, with an AI backlog reported near 73 billion dollars across hyperscaler programs including OpenAI's inference chip.[^194] TSMC wins on any branch of the tree, since GPUs, TPUs, and every competing ASIC are manufactured in the same fabs, and leading edge capacity remains the binding scarcity of the entire industry.[^278] Nvidia itself belongs on this list with a qualification the market resists: its annual cadence, Rubin in late 2026, Rubin Ultra in 2027, Feynman in 2028, is the very mechanism depreciating its customers' fleets,[^229] so Nvidia prospers as long as someone funds the treadmill, while its customers' returns are the treadmill's cost. On the device side, Apple, Qualcomm, Arm, MediaTek, and Samsung own the NPU installed base heading toward half of all consumer hardware shipments.[^197] On physical scarcity, the winners are the holders of granted grid interconnection and the suppliers of the electrical bottleneck: with transformer lead times above 160 weeks and the physical electrical layer named as the gating factor of the entire buildout as of May 2026, grid equipment order books are the scarcest asset in the chain.[^228] On demand relationships, the frontier labs with power user franchises, Anthropic in coding and agents with roughly 80 percent enterprise API revenue per third party trackers, OpenAI with 900 million weekly consumers, hold positions that a hardware repricing does not touch and cheaper serving actively improves.[^214]

### Threatened incumbents

The threatened class shares one balance sheet feature: returns that require the current paradigm, centralized, GPU bound, scarcity priced, to persist through the late 2020s. Merchant GPU neoclouds without a specialization strategy face margin compression from both directions, ASIC serving above and hardware depreciation below, with their debt collateralized by the depreciating asset itself.[^235] The late vintage mega projects of Section 11, the 2028 and 2029 delivery cohort financed at peak assumptions, carry the railway subscriber's position: committed capital, deferred delivery, paradigm risk at completion.[^240] Generic powered shell developers in weak locations hold assets whose only bid was the boom.[^243] Leveraged AI infrastructure funds hold the Situational Awareness position, and July 2026 demonstrated the exit dynamics of that trade at 4x leverage.[^246] Investors who treat AI compute as one homogeneous demand curve are mispricing both of its components.[^181] And frontier subscription businesses whose unit economics require every customer on the newest premium tier are exposed to the saturation dynamics of Section 10, in a market where the fastest growing routed models already cost under a dollar per million tokens.[^210]

### Sleepers

The sleepers are the positions the current pricing regime ignores because they monetize the transition rather than the boom. Distressed infrastructure funds and secondaries buyers are positioned for the Citadel role, acquiring half built campuses, LP stakes, and power positions at the discount the unwind creates; the fiber precedent's Level 3 style acquirers compounded for two decades on assets bought at fire sale prices.[^255] Retrofit specialists, liquid cooling providers, and modular data center vendors monetize the density mismatch, since upgrading an existing shell to host newer racks is the arbitrage between the construction clock and the silicon clock.[^226] The **information technology asset disposition** (ITAD) and accelerator recycling chain is an emerging industry with a guaranteed feedstock: AI infrastructure is projected to generate up to 2.5 million tonnes of electronic waste annually by 2030, and a functioning global resale and refurbishment chain for accelerators is already forming.[^279] Inference routing and arbitrage layers hold the toll booth of the transition: OpenRouter, routing 25 trillion tokens weekly at its May 26, 2026 Series B of 113 million dollars led by CapitalG, is paid on every token regardless of which silicon serves it.[^280] Edge runtimes and compilers (the llama.cpp and ExecuTorch ecosystems, ONNX Runtime mobile, browser native inference through WebNN and WebGPU) are the software layer of the device migration.[^201] Inference specialized chip and serving companies, Groq, Cerebras, Etched, Tenstorrent, Fireworks, Together AI, Baseten, sell the destination economics.[^195] Distillation specialists and low compute training architectures (state space model labs, the RWKV lineage, Liquid AI) monetize the efficiency curve directly.[^188] And Cloudflare class edge networks, running models across hundreds of cities at sub 100 millisecond latency, own the geography of hybrid routing.[^281]

### The observable sequence

A thesis for capital allocators must state what would prove it wrong and when. Five signals would confirm this paper's thesis, in rough expected order. First, custom ASICs taking a majority of inference accelerator spend by 2027, as mid 2026 analyses project.[^194] Second, a further break in the depreciation consensus: a second hyperscaler following Amazon's February 2025 shortening, or a material accelerator write-down at a neocloud, converting Burry's off balance sheet argument into reported numbers.[^232] Third, a credit event in GPU backed debt: a neocloud default, a forced refinancing at punitive terms, or GPU collateral haircuts repricing the sector's loan books.[^235] Fourth, conversion of the announcement pipeline into cancellations at scale: the 190 GW announced resolving toward the roughly 5 GW under construction rather than the reverse, with the 30 to 50 percent delay expectation for 2026 realized and extended into 2027.[^236] Fifth, continued decoupling of tokens from GPU revenue: token volumes compounding while merchant GPU rental pricing falls, the telecom signature of usage growth on collapsing unit prices.[^257]

Three signals would refute the thesis. First, frontier training compute sustaining 4x to 5x annual growth through 2028 with the projected 1 to 3 billion dollar runs materializing on schedule, keeping training demand large enough to absorb marginal capacity.[^268] Second, an architecture shift that resets the ASIC clock, vindicating the Hardware Lottery and restoring the GPU's flexibility premium for serving.[^192] Third, pre leasing at investment grade covenants extending to the 2028 and 2029 delivery cohort, demonstrating that demand visibility, and paying tenants rather than extrapolations, underwrite even the late vintages.[^252] Each signal in both lists is quarterly observable in public filings, MLPerf cycles, interconnection queues, and token telemetry. The thesis is falsifiable on a two year horizon.

### The phoenix scenario

Carlota Perez's surge framework, the spine of Section 2, describes technological revolutions as a sequence: installation financed by speculative capital, a frenzy, a crash at the turning point, then deployment and a golden age led by productive capital, with financial capital acting, in her words, as "the agent of massive creative destruction."[^282] The 2026 configuration matches the template's turning point markers with unusual precision: a frenzy phase running at 750 billion dollars of annual capex and 190 GW of announcements,[^236] a leveraged blowup at the financial periphery in July 2026,[^249] and an adoption curve, 7x token growth year over year at Google, that accelerated straight through the financial stress.[^218]

The phoenix scenario, graded B as a forward looking claim whose first half is observable and whose second half is not,[^283] runs as follows. The exposed vintages fail over 2027 and 2028 as delivery meets repriced serving economics; the failure is reported worldwide as the AI bubble bursting; the assets change hands at discounts, sites, megawatts, interconnection rights, and shells flowing to distressed buyers, while the silicon inside is cascaded, resold, or recycled. Serving capacity then gets rebuilt inside the acquired perimeter on transition era economics, ASIC pods, liquid cooled retrofits, edge distribution, at a cost per token far below the 2025 cost structure. Every application whose unit economics failed at 2025 serving prices becomes viable, the demand frontier of Section 6 absorbs the cheaper intelligence, and deployment accelerates on the bubble's physical leftovers, as broadband and cloud deployed on dark fiber after 2003.[^258] The honest qualifier from Section 11 applies throughout: the reusable legacy this time is the durable perimeter rather than the chips, so the golden age's bargain inputs are sites and power rather than compute itself.[^259] AI usage, on this scenario, exits the crash larger than it entered, which is what the usage data did in every prior instance of the pattern.[^257]

### Conclusion

This paper opened by refusing a phrase. "The AI bubble" bundles nine layers with nine risk profiles into one word, and the bundle fails every analytical test applied to it across thirteen sections. Separated into layers, the evidence of mid 2026 supports four conclusions. First, intelligence demand is real, measured, and compounding: 3.2 quadrillion tokens a month at one company, 900 million weekly users at another, enterprise revenue doubling and redoubling on annual timescales.[^218] Second, the wrapper layer inflates and purges continuously, release by release, and its excess deflates without accumulating into a systemic event. Third, the model layer's valuations are aggressive bets on a utility scale spending category whose early trajectory is consistent with the bet, though the big market delusion guarantees dispersion among the tickers.[^270] Fourth, the centralized compute buildout, and specifically its late, leveraged, GPU bound, weakly sited vintages, prices in the permanence of a hardware paradigm that specialization, saturation, and distribution are dismantling on an observable schedule, against construction and grid clocks that cannot keep up with an annual silicon cadence.[^229]

There is no AI bubble. There is a data center bubble, and the distinction is the most consequential capital allocation question of the coming three years. When the exposed vintages crack, the headlines will announce the end of the AI era, as the 2001 headlines announced the end of the internet era while traffic doubled through the bust.[^257] The allocator who has read the layers correctly will recognize the event as a hardware paradigm transition inside an intact intelligence boom, will be positioned in the durable winners and the sleepers before the repricing, and will be buying sites, megawatts, and interconnection rights from forced sellers at the turning point. The first forced seller has already appeared, in the last week of July 2026, and its liquidation was absorbed by a buyer who understood the difference between a broken balance sheet and a broken technology.[^249] The evidence assembled here says that difference is the entire story.

[^1]: Qianan Wang and Zen Chen, Boom, Bubble, or Buildout? A Multi-Method Evaluation of Whether Artificial Intelligence Is in an Ongoing Financial Bubble, June 2026, https://arxiv.org/abs/2606.01575
[^2]: CNBC, Hyperscalers face higher capex scrutiny after Alphabet report panned, July 28, 2026, https://www.cnbc.com/2026/07/28/hyperscalers-face-higher-capex-scrutiny-after-alphabet-report-panned.html
[^3]: CNBC, Why Leopold Aschenbrenner's Situational Awareness hedge fund imploded, July 31, 2026, https://www.cnbc.com/2026/07/31/why-leopold-aschenbrenner-situational-awareness-hedge-fund-imploded.html
[^4]: CNBC, Leopold Aschenbrenner's Situational Awareness fund fire sale, July 31, 2026, https://www.cnbc.com/2026/07/31/leopold-aschenbrenner-situational-awareness-fund-fire-sale.html
[^5]: Google, I/O keynote (Sundar Pichai), May 2025, https://blog.google/technology/ai/io-2025-keynote/ and Google I/O 2026 keynote of May 20, 2026, reported at https://blog.google
[^6]: NPR, Is an AI bubble brewing? Shiller PE Ratio nears levels seen before dot-com crash, November 13, 2025, https://www.npr.org/2025/11/13/nx-s1-5604845/is-an-ai-bubble-brewing-shiller-pe-ratio-nears-levels-seen-before-dot-com-crash
[^7]: TechTimes, AI Boom Outgrows Every Tech Bubble in History, BIS Systemic Risk Study Finds, July 15, 2026, https://www.techtimes.com/articles/320580/20260715/ai-boom-outgrows-every-tech-bubble-history-bis-systemic-risk-study-finds.htm
[^8]: Columbia University blogs, Circular financing in the AI economy, 2026, https://blogs.cuit.columbia.edu/gjb2124/circular-financing/
[^9]: Tom's Hardware, Big Tech's AI spending plans reach 725 billion dollars, 2026, https://www.tomshardware.com/tech-industry/big-tech/big-techs-ai-spending-plans-reach-725-billion
[^10]: Bloomberg, AI Circular Deals: How Microsoft, OpenAI and Nvidia Keep Paying Each Other, 2026, https://www.bloomberg.com/graphics/2026-ai-circular-deals/
[^11]: Youngjin Yoo, Ola Henfridsson, and Kalle Lyytinen, The New Organizing Logic of Digital Innovation: An Agenda for Information Systems Research, Information Systems Research 21(4), 2010, https://pubsonline.informs.org/doi/10.1287/isre.1100.0322
[^12]: Competition between AI foundation models: dynamics and policy recommendations, Industrial and Corporate Change 34(5), 2025, https://academic.oup.com/icc/article/34/5/1085/7942098
[^13]: CCIA, Intense Competition Across the AI Stack, https://ccianet.org/articles/intense-competition-across-the-ai-stack/
[^14]: ClearTake, AI Bubble Index: How to Read the Layer Gauges, 2026, https://cleartake.ai/bubble-index/
[^15]: Andrew Odlyzko, Telecom collapse (ISSCC talk and related analyses), University of Minnesota, https://www-users.cse.umn.edu/~odlyzko/talks/isscc-telecom-crash.pdf
[^16]: Paul Kedrosky, Sumant Wahi, and Alex Preston, The AI Bubble: Hidden Risks and Opportunities, Man Group, February 19, 2026, https://www.man.com/insights/the-ai-bubble
[^17]: Robert J. Shiller, Narrative Economics: How Stories Go Viral and Drive Major Economic Events, Princeton University Press, 2019, https://www.jstor.org/stable/j.ctvdf0jm5
[^18]: Narrative Emotions and Market Crises, Journal of Behavioral Finance, 2024, https://www.tandfonline.com/doi/full/10.1080/15427560.2024.2365723
[^19]: Carlota Perez, Technological Revolutions and Financial Capital: The Dynamics of Bubbles and Golden Ages, Edward Elgar, 2002, https://en.wikipedia.org/wiki/Technological_Revolutions_and_Financial_Capital
[^20]: Closelook, Technological Revolutions and Financial Capital (reading), July 5, 2026, https://www.closelook.net/library/technological-revolutions-and-financial-capital/
[^21]: Gareth Campbell, Deriving the railway mania, Financial History Review, Cambridge University Press, https://www.cambridge.org/core/journals/financial-history-review/article/abs/deriving-the-railway-mania/2EA377886FE48B4BF2DD620D4F845991
[^22]: Peter Garber, Famous First Bubbles: The Fundamentals of Early Manias, MIT Press, 2000, discussed at https://en.wikipedia.org/wiki/Tulip_mania
[^23]: Andrew Odlyzko, Collective Hallucinations and Inefficient Markets: The British Railway Mania of the 1840s, University of Minnesota, 2010, https://www.researchgate.net/publication/228291464_Collective_Hallucinations_and_Inefficient_Markets_The_British_Railway_Mania_of_the_1840s
[^24]: Gareth Campbell, The railway mania: not so great expectations?, CEPR VoxEU, https://cepr.org/voxeu/columns/railway-mania-not-so-great-expectations
[^25]: Paul A. David, The Dynamo and the Computer: A Historical Perspective on the Modern Productivity Paradox, American Economic Review 80(2), 1990, pages 355 to 361, discussed at https://en.wikipedia.org/wiki/Productivity_paradox
[^26]: Timothy Bresnahan and Manuel Trajtenberg, General Purpose Technologies: Engines of Growth?, Journal of Econometrics 65(1), 1995, https://www.nber.org/papers/w4148
[^27]: ISE Magazine, The Perils of Irrational Exuberance: The 25th Anniversary of The Dot-Com Boom, https://www.isemag.com/professional-development-leadership/article/55293922/the-perils-of-irrational-exuberancethe-25th-anniversary-of-the-dot-com-boom
[^28]: Wikipedia, Telecoms crash, https://en.wikipedia.org/wiki/Telecoms_crash
[^29]: Communications of the ACM, Dark Fiber Is Lighting Up, https://cacm.acm.org/news/dark-fiber-is-lighting-p/
[^30]: Andrew Odlyzko, Internet traffic growth: Sources and implications, Proc. SPIE ITCom, 2003, https://www-users.cse.umn.edu/~odlyzko/doc/itcom.internet.growth.pdf
[^31]: Lime, Overbuilding the Future: Why AI Is Repeating the Railways and Dark Fiber, 2026, https://lime.co/news/overbuilding-the-future-why-ai-is-repeating-the-railways-and-dark-fiber-144565/
[^32]: CNBC, The question everyone in AI is asking: How long before a GPU depreciates?, November 14, 2025, https://www.cnbc.com/2025/11/14/ai-gpu-depreciation-coreweave-nvidia-michael-burry.html
[^33]: Dave Friedman, GPU Obsolescence is Complicated, https://davefriedman.substack.com/p/gpu-obsolescence-is-complicated
[^34]: Build Inc., AI infrastructure capex and data center development, 2026, https://build.inc/insights/ai-infrastructure-capex-data-center-development
[^35]: CNBC, Leopold Aschenbrenner's hedge fund is facing steep AI losses, July 30, 2026, https://www.cnbc.com/2026/07/30/leopold-aschenbrenners-hedge-fund-is-facing-steep-ai-losses.html and CNBC, Why Leopold Aschenbrenner's Situational Awareness hedge fund imploded, July 31, 2026, https://www.cnbc.com/2026/07/31/why-leopold-aschenbrenner-situational-awareness-hedge-fund-imploded.html
[^36]: KKR, Beyond the Bubble: Why AI Infrastructure Will Compound Long after the Hype, https://www.kkr.com/insights/ai-infrastructure
[^37]: Arvind Narayanan and Akash Kapur, Up the Stack: How AI's Escape From the Commodity Trap Risks Enterprise Lock-in, Normal Technology, July 9, 2026, https://www.normaltech.ai/p/up-the-stack-how-ais-escape-from
[^38]: Labatt Simon, The AI Stack: A Map of Who Powers Enterprise AI in 2026, https://www.labattsimon.com/blog/the-ai-stack-a-map-of-who-powers-enterprise-ai-in-2026
[^39]: zylos.ai, AI chip hardware acceleration in 2026, February 1, 2026, https://zylos.ai/research/2026-02-01-ai-chip-hardware-acceleration-2026/
[^40]: The AI Consulting Network, US data centers 2026: half delayed or canceled, 2026, https://www.theaiconsultingnetwork.com/blog/us-data-centers-2026-half-delayed-canceled-cre-investors
[^41]: Mural et al., AI, Data Centers, and the U.S. Electric Grid: A Watershed Moment, Harvard Kennedy School Belfer Center, February 2026, https://www.belfercenter.org/sites/default/files/2026-02/Mural%20et%20al_AI%20Data%20Centers%20Grid_20260206.pdf
[^42]: Connect CRE, Popping the Data Center Bubble Concern, 2026, https://www.connectcre.com/stories/popping-the-data-center-bubble-concern/
[^43]: Enlit World, Data centre assets left stranded by power constraints, late 2025, https://www.enlit.world/library/data-centre-assets-left-stranded-by-power-constraints
[^44]: Utility Dive, Utilities are spending billions on the data center boom. What are the risks?, 2026, https://www.utilitydive.com/news/utilities-data-center-risk-buildout-bubble-ai-infrastructure-grid/812781/
[^45]: ValueAddVC, The Full AI Landscape 2026, https://valueaddvc.com/ai-landscape
[^46]: Martin Casado and Peter Lauten, The Empty Promise of Data Moats, Andreessen Horowitz, 2019, https://a16z.com/the-empty-promise-of-data-moats/
[^47]: Thomas Eisenmann, Geoffrey Parker, and Marshall Van Alstyne, Platform Envelopment, Harvard Business School Working Paper 07-104, 2010, https://www.hbs.edu/ris/Publication%20Files/07-104.pdf
[^48]: Joseph A. Schumpeter, Capitalism, Socialism and Democracy, 1942, creative destruction chapters, https://en.wikipedia.org/wiki/Creative_destruction
[^49]: Eric von Hippel, Democratizing Innovation, MIT Press, 2005, https://web.mit.edu/evhippel/www/democ1.htm and https://en.wikipedia.org/wiki/Eric_von_Hippel
[^50]: Nick Srnicek, Platform Capitalism, Polity, 2016, https://en.wikipedia.org/wiki/Nick_Srnicek
[^51]: David Cahn, AI's 600B Dollar Question, Sequoia Capital, June 20, 2024, https://www.sequoiacap.com/article/ais-600b-question/ and Columbia University blogs, Circular financing in the AI economy, 2026, https://blogs.cuit.columbia.edu/gjb2124/circular-financing/
[^52]: Jerker Denrell, Selection Bias and the Perils of Benchmarking, Harvard Business Review, April 2005, https://hbr.org/2005/04/selection-bias-and-the-perils-of-benchmarking
[^53]: Sebastian Mallaby, The Power Law: Venture Capital and the Making of the New Future, 2022, https://en.wikipedia.org/wiki/Sebastian_Mallaby
[^54]: Colin Camerer and Dan Lovallo, Overconfidence and Excess Entry: An Experimental Approach, American Economic Review 89(1), 1999, pages 306 to 318, https://www.aeaweb.org/articles?id=10.1257/aer.89.1.306
[^55]: core.cz, Edge computing and AI inference in 2026, https://core.cz/en/blog/2026/edge-computing-ai-inference-2026/
[^56]: DSA Research, Edge Inference vs. Centralized GPU Clusters, https://dsa-research.org/blog/edge-inference-vs-centralized-gpu/
[^57]: Blake Alcott, Jevons' paradox, Ecological Economics 54(1), 2005, pages 9 to 21, with W. S. Jevons, The Coal Question, 2nd edition, Macmillan, 1866, and the Nadella and Brynjolfsson invocations of January and February 2025, all via https://en.wikipedia.org/wiki/Jevons_paradox
[^58]: Thomas Eisenmann, Geoffrey Parker, Marshall Van Alstyne, Platform Envelopment, Harvard Business School Working Paper 07-104, revised July 27, 2010, https://www.hbs.edu/ris/Publication%20Files/07-104.pdf
[^59]: Wikipedia, Sherlock (software), https://en.wikipedia.org/wiki/Sherlock_(software)
[^60]: Arvind Narayanan and Akash Kapur, Up the Stack: How AI's Escape From the Commodity Trap Risks Enterprise Lock-in, Normal Technology, July 9, 2026, https://www.normaltech.ai/p/up-the-stack-how-ais-escape-from
[^61]: ValueAddVC, The Full AI Landscape 2026, 2026, https://valueaddvc.com/ai-landscape
[^62]: Joseph A. Schumpeter, Capitalism, Socialism and Democracy, 1942, via Wikipedia, Creative destruction, https://en.wikipedia.org/wiki/Creative_destruction
[^63]: Wikipedia, Terra (blockchain), https://en.wikipedia.org/wiki/Terra_(blockchain)
[^64]: Wikipedia, Bankruptcy of FTX, https://en.wikipedia.org/wiki/Bankruptcy_of_FTX
[^65]: Sebastian Mallaby, The Power Law: Venture Capital and the Making of the New Future, 2022, via Wikipedia, Sebastian Mallaby, https://en.wikipedia.org/wiki/Sebastian_Mallaby
[^66]: TechTimes, AI Boom Outgrows Every Tech Bubble in History, BIS Systemic Risk Study Finds, July 15, 2026, https://www.techtimes.com/articles/320580/20260715/ai-boom-outgrows-every-tech-bubble-history-bis-systemic-risk-study-finds.htm
[^67]: Eric von Hippel, Democratizing Innovation, MIT Press, 2005, https://web.mit.edu/evhippel/www/democ1.htm; lead user findings summarized via Wikipedia, Eric von Hippel, https://en.wikipedia.org/wiki/Eric_von_Hippel
[^68]: Martin Casado and Peter Lauten, The Empty Promise of Data Moats, Andreessen Horowitz, 2019, https://a16z.com/the-empty-promise-of-data-moats/
[^69]: Nick Srnicek, Platform Capitalism, Polity, 2016, via Wikipedia, Nick Srnicek, https://en.wikipedia.org/wiki/Nick_Srnicek
[^70]: David Cahn, AI's 600B Dollar Question, Sequoia Capital, June 20, 2024, https://www.sequoiacap.com/article/ais-600b-question/
[^71]: Columbia University, Circular Financing analysis, 2026, https://blogs.cuit.columbia.edu/gjb2124/circular-financing/
[^72]: Simon Labatt, The AI Stack: A Map of Who Powers Enterprise AI in 2026, 2026, https://www.labattsimon.com/blog/the-ai-stack-a-map-of-who-powers-enterprise-ai-in-2026
[^73]: Bloomberg, AI Circular Deals: How Microsoft, OpenAI and Nvidia Keep Paying Each Other, 2026, https://www.bloomberg.com/graphics/2026-ai-circular-deals/
[^74]: Colin Camerer and Dan Lovallo, Overconfidence and Excess Entry: An Experimental Approach, American Economic Review 89(1), 1999, pages 306 to 318, https://www.aeaweb.org/articles?id=10.1257/aer.89.1.306
[^75]: Jerker Denrell, Selection Bias and the Perils of Benchmarking, Harvard Business Review, April 2005, https://hbr.org/2005/04/selection-bias-and-the-perils-of-benchmarking
[^76]: io fund, AI Token Demand Is Shattering Forecasts, July 2026, https://io-fund.com/ai-stocks/ai-token-demand-shattering-forecasts
[^77]: Tuan Anh Do, The Jevons Paradox in AI Infrastructure: Efficiency, Rebound, and the Limits of Cost Reduction, SSRN working paper, 2026, https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6775299
[^78]: pdpspectra, AI Token Pricing Economics 2026, 2026, https://pdpspectra.com/blog/ai-token-pricing-economics-2026/
[^79]: io fund, Token Growth Surging: Beneficiaries, 2026, https://io-fund.com/ai-stocks/token-growth-surging-beneficiaries
[^80]: Goldman Sachs Research, AI Agents Forecast to Boost Tech Cash Flow as Usage Soars, 2026, https://www.goldmansachs.com/insights/articles/ai-agents-forecast-to-boost-tech-cash-flow-as-usage-soars
[^81]: Tuan Anh Do, The Cost Shift: Jevons' Paradox and the Restructuring of Enterprise AI Spend from Labor to Compute, SSRN working paper, 2026, https://papers.ssrn.com/sol3/papers.cfm?abstract_id=7141759
[^82]: Daron Acemoglu and Pascual Restrepo, Automation and New Tasks: How Technology Displaces and Reinstates Labor, Journal of Economic Perspectives 33(2), 2019, https://www.aeaweb.org/articles?id=10.1257%2Fjep.33.2.3
[^83]: Iain M. Cockburn, Rebecca M. Henderson, Scott Stern, The Impact of Artificial Intelligence on Innovation: An Exploratory Analysis, NBER Working Paper 24449, 2018, https://www.nber.org/papers/w24449
[^84]: NBER, When Does Automating AI Research Produce Explosive Growth? Feedback Loops in Innovation Networks, NBER Working Paper 35155, 2025, https://www.nber.org/papers/w35155
[^85]: Daron Acemoglu, The Simple Macroeconomics of AI, NBER Working Paper 32487, 2024, published in Economic Policy, 2025, https://www.nber.org/papers/w32487
[^86]: Fabrizio Dell'Acqua, Edward McFowland III, Ethan Mollick, Hila Lifshitz-Assaf, Katherine Kellogg, Saran Rajendran, Lisa Krayer, François Candelon, Karim Lakhani, Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality, Organization Science, 2025, https://pubsonline.informs.org/doi/10.1287/orsc.2025.21838
[^87]: Shakked Noy and Whitney Zhang, Experimental evidence on the productivity effects of generative artificial intelligence, Science, 2023, https://www.science.org/doi/10.1126/science.adh2586
[^88]: Zheyuan Cui, Mert Demirer, Sonia Jaffe, Leon Musolff, Sida Peng, Tobias Salz, The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers, Management Science, 2025, https://pubsonline.informs.org/doi/10.1287/mnsc.2025.00535
[^89]: PYMNTS, ChatGPT Approaches 1 Billion Weekly Active User Milestone, 2026, https://www.pymnts.com/news/artificial-intelligence/2026/chatgpt-approaches-1-billion-weekly-active-user-milestone/
[^90]: DemandSage, ChatGPT Statistics, 2026, https://www.demandsage.com/chatgpt-statistics/
[^91]: Bloomberg, OpenAI Unveils ChatGPT Work Agent to Field Tasks for Hours, July 9, 2026, https://www.bloomberg.com/news/articles/2026-07-09/openai-unveils-chatgpt-work-agent-to-field-tasks-for-hours
[^92]: Timothy F. Bresnahan and Manuel Trajtenberg, General Purpose Technologies: Engines of Growth?, Journal of Econometrics 65, 1995 (NBER Working Paper 4148, 1992), https://www.nber.org/papers/w4148
[^93]: Paul A. David, The Dynamo and the Computer: An Historical Perspective on the Modern Productivity Paradox, American Economic Review 80(2), 1990, http://www.dklevine.com/archive/refs4115.pdf
[^94]: SQ Magazine, AI Market Statistics, 2026, https://sqmagazine.co.uk/ai-market-statistics/
[^95]: AI Statistics Center, AI Market Size, 2026, https://aistatisticscenter.com/statistics/ai-market-size
[^96]: Grand View Research, Artificial Intelligence Market Analysis, https://www.grandviewresearch.com/industry-analysis/artificial-intelligence-ai-market
[^97]: MarketsandMarkets, Artificial Intelligence Market Report, https://www.marketsandmarkets.com/Market-Reports/artificial-intelligence-market-74851580.html
[^98]: TechXplore, report on peer reviewed work showing AI assistants accelerate discovery by designing and interpreting experiments, May 2026, https://techxplore.com/news/2026-05-ai-scientific-discoveries.html
[^99]: Stanford HAI, How AI Is Transforming Scientific Discovery While Keeping Humans at the Center, 2026, https://hai.stanford.edu/news/how-ai-is-transforming-scientific-discovery-while-keeping-humans-at-the-center
[^100]: International AI Safety Report 2026, https://arxiv.org/pdf/2602.21012
[^101]: Timothy F. Bresnahan and Manuel Trajtenberg, "General Purpose Technologies: Engines of Growth?", Journal of Econometrics 65, pp. 83 to 108, 1995 (NBER Working Paper 4148, 1992). https://www.nber.org/papers/w4148
[^102]: Daron Acemoglu, "The Simple Macroeconomics of AI", NBER Working Paper 32487, 2024, published in Economic Policy, 2025. https://www.nber.org/papers/w32487
[^103]: Paul A. David, "The Dynamo and the Computer: An Historical Perspective on the Modern Productivity Paradox", American Economic Review 80(2), 1990. http://www.dklevine.com/archive/refs4115.pdf
[^104]: Bloomberg, "OpenAI Unveils ChatGPT Work Agent to Field Tasks for Hours", July 9, 2026. https://www.bloomberg.com/news/articles/2026-07-09/openai-unveils-chatgpt-work-agent-to-field-tasks-for-hours
[^105]: Kerosene industrial history, documented record of United States petroleum end uses, 1859 to 1909. https://en.wikipedia.org/wiki/Kerosene
[^106]: Daniel Yergin, "The Prize: The Epic Quest for Oil, Money, and Power", Simon and Schuster, 1990. https://en.wikipedia.org/wiki/The_Prize:_The_Epic_Quest_for_Oil,_Money,_and_Power
[^107]: DemandSage, "ChatGPT Statistics", 2026. https://www.demandsage.com/chatgpt-statistics/
[^108]: ChatGPT usage statistics compilation of OpenAI usage research, July 2026. https://fatjoe.com/blog/chatgpt-stats/
[^109]: AI Statistics Center, worldwide artificial intelligence market size and spending forecast, 2026. https://aistatisticscenter.com/statistics/ai-market-size
[^110]: "Boom, Bubble, or Buildout? A Multi-Method Evaluation of Whether Artificial Intelligence Is in an Ongoing Financial Bubble", arXiv, June 2026. https://arxiv.org/pdf/2606.01575
[^111]: Carlota Perez, "Technological Revolutions and Financial Capital: The Dynamics of Bubbles and Golden Ages", Edward Elgar, 2002. https://en.wikipedia.org/wiki/Technological_Revolutions_and_Financial_Capital
[^112]: John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov et al., "Highly accurate protein structure prediction with AlphaFold", Nature 596(7873), pp. 583 to 589, 2021, DOI 10.1038/s41586-021-03819-2. https://pubmed.ncbi.nlm.nih.gov/34265844/
[^113]: Amil Merchant et al., "Scaling deep learning for materials discovery", Nature, 2023, DOI 10.1038/s41586-023-06735-9. https://pubmed.ncbi.nlm.nih.gov/38030720/
[^114]: Thomas Hubert et al., "Olympiad-level formal mathematical reasoning with reinforcement learning", Nature, 2026, DOI 10.1038/s41586-025-09833-y. https://pubmed.ncbi.nlm.nih.gov/41225005/
[^115]: TechXplore, coverage of Ali Essam Ghareeb et al., "A multi-agent system for automating scientific discovery", Nature 2026, DOI 10.1038/s41586-026-10652-y, and Juraj Gottweis et al., "Accelerating scientific discovery with Co-Scientist", Nature 2026, DOI 10.1038/s41586-026-10644-y, May 20, 2026. https://techxplore.com/news/2026-05-ai-scientific-discoveries.html
[^116]: Stanford Institute for Human-Centered Artificial Intelligence, "How AI Is Transforming Scientific Discovery While Keeping Humans at the Center", May 27, 2026. https://hai.stanford.edu/news/how-ai-is-transforming-scientific-discovery-while-keeping-humans-at-the-center
[^117]: Philip Trammell and Anton Korinek, "Economic Growth under Transformative AI", NBER Working Paper 31815, 2023, revised April 2026. https://www.nber.org/papers/w31815
[^118]: Aswath Damodaran, "SpaceX, OpenAI and Anthropic: The S&P 500 Inclusion Question and Investment Consequences", Musings on Markets, June 18, 2026. https://aswathdamodaran.blogspot.com/2026/06/
[^119]: Bradford Cornell and Aswath Damodaran, "The Big Market Delusion: Valuation and Investment Implications", Financial Analysts Journal 76(2), 2020, DOI 10.1080/0015198X.2020.1730655. https://doi.org/10.1080/0015198X.2020.1730655
[^120]: Sacra, Anthropic company profile, 2026. https://sacra.com/c/anthropic/
[^121]: SQ Magazine, "AI Market Statistics", 2026, including McKinsey State of AI survey figures, November 2025. https://sqmagazine.co.uk/ai-market-statistics/
[^122]: Sacra, OpenAI company profile, 2026. https://sacra.com/c/openai/
[^123]: Technology Checker, "ChatGPT Statistics", April 2026. https://technologychecker.io/blog/chatgpt-statistics
[^124]: Sacra, xAI company profile, 2026. https://sacra.com/c/xai/
[^125]: David Nguyen, Erik Brynjolfsson, Sophia Kazinnik, Avinash Collis and Felix Eggers, "What is Generative AI Worth?", SSRN, posted April 14, 2026, DOI 10.2139/ssrn.6569938. https://doi.org/10.2139/ssrn.6569938
[^126]: doxoINSIGHTS, "2026 U.S. Household Bill Pay Report", 2026. https://www.doxo.com/insights/2026-us-household-bill-pay-report/
[^127]: Anthropic, Claude pricing page, August 2026. https://claude.com/pricing
[^128]: Limelight Digital, "ChatGPT Users", first quarter 2026 figures. https://www.limelightdigital.co.uk/chatgpt-users/
[^129]: Telecommunications industry revenue projections (Insight Research), 2015 to 2019. https://en.wikipedia.org/wiki/Telecommunications_industry
[^130]: Ericsson, "Ericsson Mobility Report", June 2026. https://www.ericsson.com/en/reports-and-papers/mobility-report
[^131]: Grey Journal, reporting the Ramp AI Index, June 2026. https://greyjournal.net/hustle/work-tech/how-much-companies-spend-ai-tokens-2026/
[^132]: Microsoft, Microsoft 365 Copilot for business pricing, August 2026. https://www.microsoft.com/en-us/microsoft-365/copilot/business
[^133]: Menlo Ventures, "Menlo's Investment in Fireworks: The Runtime for Specialized Intelligence", July 16, 2026. https://menlovc.com/perspective/menlos-investment-in-fireworks-the-runtime-for-specialized-intelligence/
[^134]: Menlo Ventures, "2025: The State of Generative AI in the Enterprise", December 9, 2025. https://menlovc.com/perspective/2025-the-state-of-generative-ai-in-the-enterprise/
[^135]: Eleanor Wiske Dillon, Sonia Jaffe, Sida Peng and Alexia Cambon, "Early Impacts of M365 Copilot", arXiv:2504.11443, 2025. https://arxiv.org/abs/2504.11443
[^136]: Zheyuan Cui, Mert Demirer, Sonia Jaffe, Leon Musolff, Sida Peng and Tobias Salz, "The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers", Management Science, 2025. https://pubsonline.informs.org/doi/10.1287/mnsc.2025.00535
[^137]: CNBC, "Hyperscalers face higher capex scrutiny after Alphabet report panned", July 28, 2026. https://www.cnbc.com/2026/07/28/hyperscalers-face-higher-capex-scrutiny-after-alphabet-report-panned.html
[^138]: Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu and Dario Amodei, "Scaling Laws for Neural Language Models", OpenAI, arXiv, 2020. https://arxiv.org/abs/2001.08361
[^139]: Jordan Hoffmann et al., "Training Compute-Optimal Large Language Models", arXiv and NeurIPS, 2022. https://arxiv.org/abs/2203.15556
[^140]: Dario Amodei and Danny Hernandez, "AI and Compute", OpenAI, 2018. https://openai.com/index/ai-and-compute/
[^141]: Jaime Sevilla et al., "Compute Trends Across Three Eras of Machine Learning", arXiv, 2022. https://arxiv.org/pdf/2202.05924
[^142]: Epoch AI, "Training compute of frontier AI models grows by 4-5x per year", 2026. https://epoch.ai/publications/training-compute-of-frontier-ai-models-grows-by-4-5x-per-year
[^143]: Epoch AI, "Cost trend of large scale training runs", data insight, 2026. https://epoch.ai/data-insights/cost-trend-large-scale
[^144]: Epoch AI, "Model counts above compute thresholds", 2026. https://epoch.ai/publications/model-counts-compute-thresholds
[^145]: ValueAdd VC, "AI Hyperscaler Capex Compared: Why Microsoft, Google, Meta and Amazon Are All Spending at Once", 2026. https://valueaddvc.com/blog/ai-hyperscaler-capex-compared-why-microsoft-google-meta-and-amazon-are-all-spending-at-once
[^146]: DeepSeek-AI, "DeepSeek-V3 Technical Report", arXiv, 2024. https://arxiv.org/abs/2412.19437
[^147]: Michael C. Frank, "Bridging the data gap between children and large language models", Trends in Cognitive Sciences, 2023. https://www.sciencedirect.com/science/article/abs/pii/S1364661323002036
[^148]: Pablo Villalobos, Anson Ho, Jaime Sevilla, Tamay Besiroglu, Lennart Heim and Marius Hobbhahn, "Will we run out of data? Limits of LLM scaling based on human-generated data", Epoch AI, 2022, updated 2024, ICML 2024. https://epoch.ai/publications/will-we-run-out-of-data-limits-of-llm-scaling-based-on-human-generated-data
[^149]: Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Yarin Gal, Nicolas Papernot and Ross Anderson, "The Curse of Recursion: Training on Generated Data Makes Models Forget", arXiv, 2023, Nature version 2024. https://arxiv.org/abs/2305.17493
[^150]: FuturePicker, "AI Training Data Crisis and Synthetic Collapse", 2026. https://futurepicker.com/en/ai-training-data-crisis-synthetic-collapse-2026-en/
[^151]: Anson Ho, Tamay Besiroglu, Ege Erdil, David Owen, Robi Rahman, Zifan Carl Guo, David Atkinson, Neil Thompson and Jaime Sevilla, "Algorithmic progress in language models", NeurIPS 2024, arXiv, 2024. https://arxiv.org/abs/2403.05812
[^152]: Danny Hernandez and Tom B. Brown, "Measuring the Algorithmic Efficiency of Neural Networks", OpenAI, arXiv, 2020. https://arxiv.org/abs/2005.04305
[^153]: Ege Erdil and Tamay Besiroglu, "Algorithmic progress in computer vision", arXiv, 2022. https://arxiv.org/pdf/2212.05153
[^154]: Epoch AI, "Algorithmic progress in language models", blog analysis including Shapley decomposition. https://epoch.ai/blog/algorithmic-progress-in-language-models
[^155]: ValueAdd VC, "How AI Inference Costs Have Dropped 95% in Two Years and What Happens Next", 2026. https://valueaddvc.com/blog/how-ai-inference-costs-have-dropped-95-in-two-years-and-what-happens-next
[^156]: AI Superior, "LLM Token Cost", 2026 pricing guide. https://aisuperior.com/llm-token-cost/
[^157]: GPUnex, "AI Inference Economics 2026", 2026. https://www.gpunex.com/blog/ai-inference-economics-2026/
[^158]: Ben Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli and Ari Morcos, "Beyond neural scaling laws: beating power law scaling via data pruning", NeurIPS 2022 Outstanding Paper. https://arxiv.org/abs/2206.14486
[^159]: DeepSeek-AI, "DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning", Nature 645, pp. 633 to 638, 2025. https://arxiv.org/abs/2501.12948
[^160]: "QUEST: Training Frontier Deep Research Agents with Fully Synthetic Tasks", arXiv, 2026. https://arxiv.org/abs/2605.24218
[^161]: Semiautonomous Systems, "The Training Data Ecosystem in 2026", 2026. https://semiautonomous.systems/blog/training-data-ecosystem-2026/
[^162]: Epoch AI, "The gap between open and closed models", data insight, May 29, 2026. https://epoch.ai/data-insights/open-closed-eci-gap
[^163]: LessWrong, "How far behind are open models?", analysis. https://www.lesswrong.com/posts/rJcCrXyEsJKmmDpWG/how-far-behind-are-open-models
[^164]: Chong Zhang et al., "100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models", arXiv, 2025. https://arxiv.org/abs/2505.00551
[^165]: Hugging Face, Open-R1 project, launched January 2025. https://github.com/huggingface/open-r1
[^166]: SoloSoft, "TinyZero and the R1 Reproduction Wave", 2026. https://www.solosoft.dev/post/tinyzero-r1-reproduction-2026/
[^167]: Tokenscost, "Hugging Face Sustainable Monetization", 2026 figures as of May 2026. https://tokenscost.com/blog/hugging-face-sustainable-monetization-2026
[^168]: AgentMarketCap, "Hugging Face Agents Hub: Open Ecosystem versus Proprietary Marketplaces", April 8, 2026. https://agentmarketcap.ai/blog/2026/04/08/hugging-face-agents-hub-open-ecosystem-vs-proprietary-marketplaces
[^169]: Programming Helper, "Hugging Face 2026: 2 million models and download distribution", 2026. https://www.programming-helper.com/tech/hugging-face-2026-2-million-models-80-percent-downloads-python
[^170]: Grogan, "The End of the Foundation Model Era: Open-Weight Models, Sovereign AI, and Inference as Infrastructure", arXiv, submitted March 18, 2026. https://arxiv.org/abs/2604.06217
[^171]: Agent Wars, "DeepSeek, Open Weights and the AI Moat", April 28, 2026. https://www.agent-wars.com/news/2026-04-28-deepseek-open-weights-ai-moat
[^172]: Amy Zegart and Emerson Johnston, "A Deep Peek into DeepSeek AI's Talent and Implications for US Innovation", Hoover Institution, April 2025. https://www.hoover.org/research/deep-peek-deepseek-ais-talent-and-implications-us-innovation
[^173]: ChinaTalk, "DeepSeek's Secret to Success", 2025. https://www.chinatalk.media/p/deepseeks-secret-to-success
[^174]: WhySoGeek, "Meta Superintelligence Labs AI Talent War", June 2026. https://whysogeek.com/meta-superintelligence-labs-ai-talent-war-june-2026/
[^175]: JobsByCulture, "AI Engineer and Researcher Salary", 2026 compensation survey. https://jobsbyculture.com/blog/ai-engineer-researcher-salary-2026
[^176]: Build Fast with AI, "LLM Scaling Laws Explained", 2026. https://www.buildfastwithai.com/blogs/llm-scaling-laws-explained
[^177]: Sara Hooker, "On the Slow Death of Scaling", SSRN, 2026. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5877662
[^178]: Albert Gu and Tri Dao, "Mamba: Linear-Time Sequence Modeling with Selective State Spaces", COLM 2024, arXiv, 2023. https://arxiv.org/abs/2312.00752
[^179]: Victor Dibia, "Is Scaling a Dead End? Why Model Scaling is Necessary Infrastructure", 2026. https://newsletter.victordibia.com/p/is-scaling-a-dead-end-why-model-scaling
[^180]: Nikhil Sardana and Jonathan Frankle, "Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling Laws", arXiv, 2024. https://arxiv.org/pdf/2401.00448
[^181]: Epoch AI (Erdil et al.), Optimally Allocating Compute Between Inference and Training, https://epoch.ai/publications/optimally-allocating-compute-between-inference-and-training; Epoch AI, Trading Off Compute in Training and Inference, https://epoch.ai/publications/trading-off-compute-in-training-and-inference
[^182]: Computerworld, CES 2026: AI compute sees a shift from training to inference (Deloitte estimate), January 2026, https://www.computerworld.com/article/4114579/ces-2026-ai-compute-sees-a-shift-from-training-to-inference.html
[^183]: Intellectia, NVDA Stock Earnings Analysis May 2026, https://intellectia.ai/blog/nvda-stock-earnings-analysis-may-2026; Nvidia earnings call, May 20, 2026
[^184]: Nvidia, press release announcing financial results, May 2026, https://investor.nvidia.com/news/press-release-details/2026/NVIDIA-Announces-Financial-Results
[^185]: AgentMarketCap, AI inference overtakes training: two thirds of 2026 compute, April 17, 2026, https://agentmarketcap.ai/blog/2026/04/17/ai-inference-overtakes-training-two-thirds-2026-compute; ValueAddVC, Inference chips vs training chips: why the next semiconductor race is different, https://valueaddvc.com/blog/inference-chips-vs-training-chips-why-the-next-semiconductor-race-is-different
[^186]: Epoch AI Brief, August 2025, https://epochai.substack.com/p/the-epoch-ai-brief-august-2025
[^187]: Ho et al., Algorithmic progress in language models, NeurIPS 2024, https://arxiv.org/abs/2403.05812
[^188]: Gundlach, Lynch, Thompson, Meek Models Shall Inherit the Earth, 2025, https://arxiv.org/abs/2507.07931
[^189]: Epoch AI, Training compute of frontier AI models grows by 4-5x per year, https://epoch.ai/publications/training-compute-of-frontier-ai-models-grows-by-4-5x-per-year
[^190]: Jouppi et al., In Datacenter Performance Analysis of a Tensor Processing Unit, ISCA 2017, https://arxiv.org/abs/1704.04760
[^191]: Jouppi et al., TPU v4: An Optically Reconfigurable Supercomputer for Machine Learning, ISCA 2023, https://arxiv.org/abs/2304.01433
[^192]: Sara Hooker, The Hardware Lottery, 2020, https://arxiv.org/abs/2009.06489
[^193]: Spheron, Google TPU v7 Ironwood vs NVIDIA B200: Inference Cost, 2026, https://www.spheron.network/blog/google-tpu-v7-ironwood-vs-nvidia-b200-inference-cost/
[^194]: Industry analyses aggregated mid 2026 (third party figures, reported with caution), https://html.duckduckgo.com/html/?q=ASIC+vs+GPU+inference+2026+custom+silicon+shift+Google+TPU+Broadcom+OpenAI+Meta+MTIA
[^195]: Vendor and comparison figures aggregated 2026 (vendor claims, reported with caution), https://html.duckduckgo.com/html/?q=Groq+LPU+energy+efficiency+cost+per+token+vs+GPU+Cerebras+inference+MLPerf+2026
[^196]: ByteIota, AI Inference Costs 2026: The Hidden 15 to 20x GPU Crisis, 2026, https://byteiota.com/ai-inference-costs-2026-the-hidden-15-20x-gpu-crisis/
[^197]: Counterpoint Research, GenAI smartphone share to rise to 45 percent of global shipments in 2026, https://www.counterpointresearch.com/en/insights/genai-smartphone-share-to-rise-to-45-percent-of-global-shipments-in-2026
[^198]: Communications Today via Counterpoint, GenAI smartphones to make up 45 percent of global shipments in 2026, https://www.communicationstoday.co.in/genai-smartphones-to-make-up-45-of-global-shipments-in-2026/
[^199]: AIPC Index, Copilot Plus share of Windows notebook shipments, June 2026, https://aipc.computer/index
[^200]: PC Central, Apple Intelligence iPhone shipments hit 450 million in Q1 2026, https://pccentral.net/apple-intelligence-iphone-shipments-hit-450-million-q1-2026/
[^201]: On Device Language Models: A Comprehensive Review, 2024, https://arxiv.org/html/2409.00088v1
[^202]: Mobile Edge Intelligence for Large Language Models, 2024, https://arxiv.org/abs/2407.18921
[^203]: SolidAITech, NPU Guide: TOPS Marketing vs Reality, June 2026, https://www.solidaitech.com/2026/06/npu-guide-tops-marketing-vs-reality.html
[^204]: Data Gate, Edge AI vs cloud cost analysis 2026, https://data-gate.ch/edge-ai-vs-cloud-cost-analysis-2026/
[^205]: Clayton M. Christensen, The Innovator's Dilemma, Harvard Business School Press, 1997; mechanism summary at https://www.christenseninstitute.org/theory/disruptive-innovation/
[^206]: Herbert A. Simon, Rational Choice and the Structure of the Environment, Psychological Review 63, 1956, https://pubmed.ncbi.nlm.nih.gov/13310708/
[^207]: SellCell, How Often Do People Upgrade Their Phone? (2026 statistics page, underlying data 2025), https://www.sellcell.com/blog/how-often-do-people-upgrade-their-phone/
[^208]: TechCrunch, Should you still buy your next smartphone or subscribe to it instead?, August 1, 2026, https://techcrunch.com/2026/08/01/should-you-still-buy-your-next-smartphone-or-subscribe-to-it-instead/
[^209]: iDropNews, How Long Do People Keep iPhones? 2026 Upgrade Data, https://www.idropnews.com/news/why-iphone-upgrade-cycles-are-longer/261990/
[^210]: Digital Applied, OpenRouter Rankings April 2026: Top AI Models by Data, https://www.digitalapplied.com/blog/openrouter-rankings-april-2026-top-ai-models-data
[^211]: Deneckere and McAfee, Damaged Goods, Journal of Economics and Management Strategy 5(2), 1996, https://onlinelibrary.wiley.com/doi/10.1111/j.1430-9134.1996.00149.x
[^212]: Publicis Sapient, 2026 Global Enterprise AI Report, https://www.publicissapient.com/company/news/ai-adoption-enterprise-readiness-report-2026
[^213]: Eric von Hippel, Lead Users: A Source of Novel Product Concepts, Management Science 32(7), 1986, https://pubsonline.informs.org/doi/10.1287/mnsc.32.7.791
[^214]: Anthropic Economic Index, https://www.anthropic.com/economic-index; third party revenue trackers, mid 2026, aggregated at https://html.duckduckgo.com/html/?q=Anthropic+revenue+2026+API+enterprise+share+ARR+coding (unconfirmed figures, reported with caution)
[^215]: ValueAddVC, OpenAI Revenue 2026, https://valueaddvc.com/blog/openai-revenue-2026-20b-arr-4b-month-path-to-profitability
[^216]: VentureBeat, Frontier models are failing one in three production attempts, and getting harder to audit, 2026, https://venturebeat.com/security/frontier-models-are-failing-one-in-three-production-attempts-and-getting-harder-to-audit
[^217]: Fidji Simo, Closing the capability gap between frontier AI and everyday use, 2026, https://fidjisimo.substack.com/p/closing-the-capability-gap
[^218]: Sundar Pichai, Google I/O 2026 keynote, May 20, 2026, https://blog.google/innovation-and-ai/sundar-pichai-io-2026/
[^219]: Nodemini, OpenRouter token weekly rankings and billing market share, week of May 18 to 24, 2026, https://nodemini.com/en/blog/2026-openrouter-token-weekly-rankings-billing-market-share.html
[^220]: AIBusinessWeekly, OpenAI statistics, July 2026, https://aibusinessweekly.net/p/openai-statistics; CNBC, OpenAI CFO Sarah Friar tells employees ARR in July topped all of Q2, July 29, 2026, https://www.cnbc.com/2026/07/29/openai-cfo-sarah-friar-tells-employees-arr-in-july-topped-all-of-q2.html
[^221]: Mobile Edge Intelligence for Large Language Models, 2024, https://arxiv.org/abs/2407.18921; Spheron, AI Inference Cost Economics 2026, https://spheron.network/blog/ai-inference-cost-economics-2026/
[^222]: Epoch AI inference forecast, reported January 2026, https://phemex.com/news/article/epoch-ai-forecasts-inference-compute-to-surpass-model-training-by-2030-86253
[^223]: Grogan, The End of the Foundation Model Era: Open-Weight Models, Sovereign AI, and Inference as Infrastructure, 2026, https://arxiv.org/abs/2604.06217
[^224]: Ben Caldecott, Introduction to special issue: stranded assets and the environment, Journal of Sustainable Finance and Investment, 2017, https://www.tandfonline.com/doi/full/10.1080/20430795.2016.1266748
[^225]: Caldecott et al., Stranded Assets: Environmental Drivers, Societal Challenges, and Supervisory Responses, Annual Review of Environment and Resources, 2021, https://www.annualreviews.org/content/journals/10.1146/annurev-environ-012220-101430
[^226]: Bill Dollins, AI Data Centers and the Risk of Stranded Infrastructure, June 15, 2026, https://blog.geomusings.com/2026/06/15/ai-data-centers-and-the-risk-of-stranded-infrastructure
[^227]: Data Center Knowledge, Gridlocked: How Power Constraints Are Shaping the Future of Data Centers, https://www.datacenterknowledge.com/energy-power-supply/gridlocked-how-power-constraints-are-shaping-the-future-of-data-centers
[^228]: ATK Energy, Grid Interconnection for Data Centers in 2026, 2026, https://atkenergygroup.com/blog/grid-interconnection-data-centers/
[^229]: Hyperframe Research, NVIDIA's Annual Roadmap Rhythm: Innovation Engine or CapEx Quicksand?, January 7, 2026, https://hyperframeresearch.com/2026/01/07/nvidias-annual-roadmap-rhythm-innovation-engine-or-capex-quicksand/; corroborated by GTC 2026 coverage, Technology Pulse, May 22, 2026, https://technologypulse.app/2026-05-22-nvidia-rubin-architecture/
[^230]: Mural et al., AI, Data Centers, and the U.S. Electric Grid: A Watershed Moment, Harvard Belfer Center, February 2026, https://www.belfercenter.org/sites/default/files/2026-02/Mural%20et%20al_AI%20Data%20Centers%20Grid_20260206.pdf
[^231]: Dave Friedman, The 176 billion dollar accounting question, https://davefriedman.substack.com/p/the-176-billion-accounting-question
[^232]: MBI Deep Dives, Big tech earnings quality, https://www.mbi-deepdives.com/big-tech-earnings-quality/; Deep Quarry, Depreciation of GPUs: between useful life and obsolescence, https://deepquarry.substack.com/p/depreciation-of-gpus-between-useful
[^233]: State Street Global Advisors, Why the AI CapEx cycle may have more staying power than you think, https://www.ssga.com/us/en/institutional/insights/ai-capex-cycle-may-have-more-staying-power
[^234]: Nvidia depreciation commentary, via Deep Quarry, https://deepquarry.substack.com/p/depreciation-of-gpus-between-useful
[^235]: Hashrate Index, Used GPU market: A100 and H100 pricing and depreciation, https://hashrateindex.com/blog/used-gpu-market-pricing-deprecation-secondary-ai/
[^236]: Build Inc., AI Infrastructure Capex Is Rewriting Data Center Development in 2026, https://build.inc/insights/ai-infrastructure-capex-data-center-development
[^237]: Tech Insider, U.S. AI Data Center Delays: 7 GW Capacity Crisis, 2026, https://tech-insider.org/us-ai-data-center-delays-cancellations-7gw-capacity-crisis-2026/
[^238]: Epoch AI, OpenAI Stargate: where the US sites stand, https://epoch.ai/publications/openai-stargate-where-the-us-sites-stand; Data Center Knowledge, Stargate Update: AI's Biggest Data Center Buildout Meets Reality, https://www.datacenterknowledge.com/ai-data-centers/stargate-update-ai-s-biggest-data-center-buildout-meets-reality; Tradeline, July 2026, https://www.tradelineinc.com/news/2026-7/oracle-and-openai-celebrate-construction-stargate-data-center
[^239]: Andrew Odlyzko, Collective Hallucinations and Inefficient Markets: The British Railway Mania of the 1840s, 2010, https://www.researchgate.net/publication/228291464_Collective_Hallucinations_and_Inefficient_Markets_The_British_Railway_Mania_of_the_1840s
[^240]: Gareth Campbell, The railway mania: not so great expectations?, CEPR VoxEU, https://cepr.org/voxeu/columns/railway-mania-not-so-great-expectations
[^241]: Andrew Odlyzko, The railway mania of the 1860s and financial innovation, https://www-users.cse.umn.edu/~odlyzko/doc/mania18.pdf
[^242]: TechTimes, AI Boom Outgrows Every Tech Bubble in History, BIS Systemic Risk Study Finds, July 15, 2026, https://www.techtimes.com/articles/320580/20260715/ai-boom-outgrows-every-tech-bubble-history-bis-systemic-risk-study-finds.htm
[^243]: Bill Dollins, AI Data Centers and the Risk of Stranded Infrastructure, June 15, 2026, https://blog.geomusings.com/2026/06/15/ai-data-centers-and-the-risk-of-stranded-infrastructure
[^244]: Barchart, Leopold Aschenbrenner's Situational Awareness surpasses 20 billion dollars as AI focused hedge fund gains 270 percent in 2026, https://www.barchart.com/story/news/2370002/leopold-aschenbrenners-situational-awareness-surpasses-20-billion-as-ai-focused-hedge-fund-gains-270-in-2026
[^245]: Disruption Banking, Can the Situational Awareness hedge fund raise capital after its 439 percent H1 gain?, July 30, 2026, https://www.disruptionbanking.com/2026/07/30/can-the-situational-awareness-hedge-fund-raise-capital-after-its-439-h1-gain/
[^246]: CNBC, Why Leopold Aschenbrenner's Situational Awareness hedge fund imploded, July 31, 2026, https://www.cnbc.com/2026/07/31/why-leopold-aschenbrenner-situational-awareness-hedge-fund-imploded.html
[^247]: CNBC, Why Leopold Aschenbrenner's Situational Awareness hedge fund imploded, July 31, 2026, https://www.cnbc.com/2026/07/31/why-leopold-aschenbrenner-situational-awareness-hedge-fund-imploded.html
[^248]: CNBC, Leopold Aschenbrenner's hedge fund is facing steep AI losses, July 30, 2026, https://www.cnbc.com/2026/07/30/leopold-aschenbrenners-hedge-fund-is-facing-steep-ai-losses.html; Bloomberg, Aschenbrenner hedge fund Situational Awareness seeks capital after loss, July 30, 2026, https://www.bloomberg.com/news/articles/2026-07-30/aschenbrenner-hedge-fund-situational-awareness-seeks-capital-after-loss-ft-says
[^249]: CNBC, Leopold Aschenbrenner Situational Awareness fund fire sale, July 31, 2026, https://www.cnbc.com/2026/07/31/leopold-aschenbrenner-situational-awareness-fund-fire-sale.html
[^250]: SpotGamma, Anatomy of a Margin Call: How Situational Awareness LP Unwound a 20 Billion Dollar AI Book in One Trade, https://spotgamma.com/situational-awareness-unwind-margin-call-ai/
[^251]: Rebellion Research, The rise, the margin, the drawdown: the lesson of Situational Awareness, https://www.rebellionresearch.com/the-rise-the-margin-the-drawdown-the-lesson-of-situational-awareness
[^252]: Connect CRE, Popping the Data Center Bubble Concern, 2026, https://www.connectcre.com/stories/popping-the-data-center-bubble-concern/; Build Inc., May 2026 assessment, https://build.inc/insights/ai-infrastructure-capex-data-center-development
[^253]: Andrew Odlyzko, Internet growth: Myth and reality, use and abuse, 2001, https://www-users.cse.umn.edu/~odlyzko/doc/internet.growth.myth2.pdf; Broadband Breakfast, Broadband expert Andrew Odlyzko warns telecom investors that industry has its math wrong again, https://broadbandbreakfast.com/broadband-expert-andrew-odlyzko-warns-telecom-investors-that-industry-has-its-math-wrong-again/
[^254]: ISE Magazine, The Perils of Irrational Exuberance: The 25th Anniversary of The Dot-Com Boom, https://www.isemag.com/professional-development-leadership/article/55293922/the-perils-of-irrational-exuberancethe-25th-anniversary-of-the-dot-com-boom
[^255]: The Bubble Bubble, The Late 1990s Telecom Bubble, https://www.thebubblebubble.com/telecom-bubble/; Fabricated Knowledge, Lessons from History: The Rise and Fall of the Telecom Bubble, https://www.fabricatedknowledge.com/p/lessons-from-history-the-rise-and
[^256]: Federal Reserve Bank of Richmond, Boom and Bust in Telecommunications, Economic Quarterly, 2003, https://www.richmondfed.org/~/media/richmondfedorg/publications/research/economic_quarterly/2003/fall/pdf/wolman.pdf
[^257]: Andrew Odlyzko, Internet traffic growth: Sources and implications, Proc. SPIE ITCom, 2003, https://www-users.cse.umn.edu/~odlyzko/doc/itcom.internet.growth.pdf
[^258]: Opus Interactive, The Telecom Overbuild That Built the Cloud: Lessons from Fiber Glut to Hyperscale Success, https://www.opusinteractive.com/opus-blog/post/the-telecom-overbuild-that-built-the-cloud-lessons-from-fiber-glut-to-hyperscale-success/; Communications of the ACM, Dark Fiber Is Lighting Up, https://cacm.acm.org/news/dark-fiber-is-lighting-p/
[^259]: Dave Friedman, GPU Obsolescence is Complicated, https://davefriedman.substack.com/p/gpu-obsolescence-is-complicated
[^260]: Andrew Odlyzko, telecom crash analyses, https://www-users.cse.umn.edu/~odlyzko/talks/isscc-telecom-crash.pdf
[^261]: Ilya Sutskever, NeurIPS 2024 remarks, via Dibia, Is Scaling a Dead End? Why Model Scaling is Necessary Infrastructure, 2026, https://newsletter.victordibia.com/p/is-scaling-a-dead-end-why-model-scaling
[^262]: The Economics of LLM Inference: Batch Economics, https://mlechner.substack.com/p/the-economics-of-llm-inference-batch
[^263]: Genius Tech Lab, Reasoning models break inference economics, June 28, 2026, https://geniustechlab.com/posts/2026-06-28-reasoning-models-test-time-compute-economics
[^264]: Blake Alcott, Jevons' paradox, Ecological Economics 54(1), 2005; W. S. Jevons, The Coal Question, 2nd ed., Macmillan, 1866; both via https://en.wikipedia.org/wiki/Jevons_paradox
[^265]: Satya Nadella's invocation of the Jevons paradox for AI, Barron's, January 28, 2025, via https://en.wikipedia.org/wiki/Jevons_paradox
[^266]: SSRN, Jevons effects in AI infrastructure, https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6775299; Goldman Sachs, AI agents forecast to boost tech cash flow as usage soars, https://www.goldmansachs.com/insights/articles/ai-agents-forecast-to-boost-tech-cash-flow-as-usage-soars
[^267]: Tesorb, Frontier AI models Q2 2026 tracker, https://tesorb.com/frontier-ai-models-q2-2026-tracker/
[^268]: Deluair, Frontier AI training cost 2026, https://deluair.com/consultancy/insights/frontier-ai-training-cost-2026; Varun K., What it takes to train a frontier LLM, https://www.varunk.me/blog/what-it-takes-to-train-a-frontier-llm
[^269]: David Cahn, AI's 600B Question, Sequoia Capital, June 2024, https://www.sequoiacap.com/article/ais-600b-question/
[^270]: Bradford Cornell and Aswath Damodaran, The Big Market Delusion: Valuation and Investment Implications, Financial Analysts Journal 76(2), 2020, DOI 10.1080/0015198X.2020.1730655
[^271]: McKinsey State of AI survey, November 2025, cited at https://sqmagazine.co.uk/ai-market-statistics/
[^272]: AIBusinessWeekly, OpenAI statistics, July 2026, https://aibusinessweekly.net/p/openai-statistics; CNBC, OpenAI CFO Sarah Friar tells employees ARR in July topped all of Q2, July 29, 2026, https://www.cnbc.com/2026/07/29/openai-cfo-sarah-friar-tells-employees-arr-in-july-topped-all-of-q2.html
[^273]: Sacra, Anthropic company profile, https://sacra.com/c/anthropic/
[^274]: Nguyen, Brynjolfsson, Kazinnik, Collis, Eggers, consumer willingness to accept study, DOI 10.2139/ssrn.6569938; doxo, 2026 US Household Bill Pay Report, https://www.doxo.com/insights/2026-us-household-bill-pay-report/
[^275]: Igor Markov, The False Dawn: Reevaluating Google's Reinforcement Learning for Chip Macro Placement, 2023, https://arxiv.org/abs/2306.09633; Communications of the ACM reevaluation, https://cacm.acm.org/research/reevaluating-googles-reinforcement-learning-for-ic-macro-placement/
[^276]: Mirhoseini, Goldie et al., A graph placement methodology for fast chip design, Nature 594, 2021, https://www.nature.com/articles/s41586-021-03544-w; Goldie, Mirhoseini, Yazgan et al., Addendum, Nature 634, October 2024, DOI 10.1038/s41586-024-08032-5
[^277]: Synopsys DSO.ai, https://www.synopsys.com/ai/ai-powered-eda/dso-ai.html; Cadence Cerebrus AI Studio, https://www.cadence.com/en_US/home/tools/digital-design-and-signoff/soc-implementation-and-floorplanning/cadence-cerebrus-ai-studio.html; DAC 2026 announcements via https://html.duckduckgo.com/html/?q=Synopsys+Cadence+AI+EDA+2026+adoption+DSO.ai+Cerebrus+generative+AI+chip+design
[^278]: China capacity reporting aggregated 2026 (third party figures, reported with caution), https://html.duckduckgo.com/html/?q=China+chip+production+capacity+2026+SMIC+expansion+Huawei+Ascend+status
[^279]: UN News, AI infrastructure e-waste projections, June 2026, https://news.un.org/en/story/2026/06/1167658; Wang, Zhang, Tzachor, Chen, E-waste challenges of generative artificial intelligence, Nature Computational Science, 2024, https://www.nature.com/articles/s43588-024-00712-6; Reloop Global, AI Hardware Recycling Guide 2026, https://reloopglobal.com/blog/ai-hardware-recycling-gpu-servers/
[^280]: Business Wire, OpenRouter Series B announcement, May 26, 2026, https://www.businesswire.com/news/home/20260526953416/en/
[^281]: Core.cz, Edge computing and AI inference 2026, https://core.cz/en/blog/2026/edge-computing-ai-inference-2026/
[^282]: Carlota Perez, Technological Revolutions and Financial Capital: The Dynamics of Bubbles and Golden Ages, Edward Elgar, 2002, via https://en.wikipedia.org/wiki/Carlota_Perez
[^283]: Closelook, Technological Revolutions and Financial Capital, July 5, 2026, https://www.closelook.net/library/technological-revolutions-and-financial-capital/
