Every publication is available in Chinese, English, and Arabic每篇内容均提供中文、英文和阿拉伯文版本

All writing

Marginal Cost Is Breaking Down: When "Cognitive Capability" Becomes a Reproducible, Ownable Factor of Production

**By PeterZou**

This essay is available in three complete language versions

**By PeterZou**

02I. A Story That Has Been Told to Death—and Still Told Wrong

For the past two years, nearly every AI product launch has repeated the same line: "Intelligence is becoming as cheap as tap water."

The data seems to support the narrative: within a few years, the inference price of large models at an equivalent level of capability has fallen by several orders of magnitude. And so "marginal cost trending to zero" became entrepreneurs' default premise—since the cost of a call will keep dropping, just pile on agents, pile on tokens, pile on automation, and the business model will naturally work itself out.

But if you have actually watched the bills, you will notice a glaring contradiction: **unit prices are plunging, while total spending is rising.**

This is not a bug. It is the single most misread aspect of this transformation. What is actually happening is not "AI makes things cheaper"—that is merely a capability upgrade, an optimization inside the old paradigm. What is actually happening is that **"cognitive capability" has, for the first time, become a factor of production that can be owned, reproduced, depreciated, and traded**—and the way marginal cost is determined has shifted from "exogenously given by technology" to "endogenously determined by capability tier and task structure."

The former is deepening. The latter is revolution. And conflating the two is the source of most strategic misjudgment today.

03II. The Test: What Actually Counts as a Paradigm Revolution

To judge whether something is a revolution, you cannot look at how much the price dropped. You have to look at whether **the definitions changed**.

A Kuhnian paradigm revolution is not parameter tuning; it is a change in the standards for "what counts as a problem and what counts as a good answer." Translated into economics, I use only two criteria:

**Criterion One: Has the "factor identity" of production changed?** (What can be owned, reproduced, depreciated, and traded?)

**Criterion Two: Has the way cost is determined changed?** (Is marginal cost an exogenous technological parameter, or an endogenous strategic variable?)

The old paradigm's default answers are clear: the production function is written as Y = F(K, L), with the capital–labor dichotomy; technology is an exogenous total factor productivity residual that firms cannot buy; marginal cost is positive and declines with output; knowledge is a non-rival public good, and patents and copyrights manufacture **artificial scarcity**【verified · classic literature】.

Note a fact that is widely overlooked: **the old paradigm never denied "zero marginal cost."** That information goods have "high fixed costs and near-zero marginal reproduction costs" was systematically written into *Information Rules* by Shapiro and Varian back in 1999【verified · classic literature】. The old paradigm's way of handling it was to place it in a box labeled "special case plus institutional patch."

So "the reproduction cost of some commodity trending to zero" is not a revolution at all. **The criterion for revolution is this: even the cost of "production capability" itself has become a variable that can be reproduced and endogenously determined.**

04III. The Evidence: The Collapse Is Stratified, and the Bill Can Move in Reverse

**First, the collapse itself: it is real, but highly stratified.**

Using MMLU-equivalent performance as the metric, a16z claims LLM inference costs fall roughly 10x per year; the minimum cost to reach an MMLU score of 42 fell from $60 per million tokens in the GPT-3 era to $0.06—a 1,000x drop in three years【verified】. Using performance milestones as the metric, Epoch AI claims the inference price to reach GPT-4-level performance falls roughly 40x per year, but the rate varies from 9x to 900x per year depending on the milestone【verified】. After re-checking with the largest benchmark-price dataset, MIT FutureTech and others reach a more conservative and more credible conclusion: the price–performance of frontier models improves by roughly 5–10x per year, and algorithm efficiency improves by roughly 3x per year after controlling for hardware progress; but the highest performance tier falls 31x per year, while the lowest tier falls only 1.7x【verified】.

The conclusion is cold: **"marginal cost trending to zero" is task-dependent and capability-stratified—it is not a universal march to zero.** What is truly collapsing is "the price of buying a given capability," while low-end capability barely moves.

**Now, why do the bills rise?**

First, the cost structure is a "dumbbell." Epoch AI estimates that frontier model training costs have grown 2.4x per year since 2016; GPT-4 training cost roughly $40 million, and on this trajectory the largest training run in 2027 will exceed $1 billion【verified】. **Fixed costs are rising exponentially while marginal costs are falling exponentially.** Economies of scale no longer show up as "falling unit product cost" but as "rising capability reuse rate"—increasing returns are concentrated in the hands of the very few players who control compute.

Second, generation got cheaper, but verification got more expensive. A team at MIT found a counterintuitive fact: despite the collapse in model price–performance, **the cost of running benchmark evaluations has stayed flat or even risen**. A single-model evaluation on SWE-bench Verified can reach thousands of dollars, and a single breakthrough evaluation on ARC-AGI reportedly costs around $3,000【verified】. The reason is that reaching higher performance requires larger models and longer reasoning chains, which directly cancel out the falling unit price.

This is precisely the migration of Baumol and Bowen's "cost disease"【inference】: as the cost of the generation sector collapses, the complementary verification, judgment, and evaluation sectors advance more slowly, and their relative prices rise instead. **The value bottleneck shifts from "production" to "verification, specification, trust, and accountability."**

Third, efficiency gains generate demand expansion. Industry analyses point out that reasoning models produce large numbers of thinking tokens billed by output, and agent loops resend accumulated context over and over; token consumption for a task can be 5–30x that of an equivalent chat task, and one ticket-handling agent costs roughly 50–60x per task compared with a chatbot【single source · pending first-hand verification】. This is the Jevons paradox【verified · classic literature + inference】: **falling prices are the cause of the consumption explosion, not a guarantee that the bill will fall.**

**Finally, a counterexample must be placed alongside: the macro level is not yet verified.** Using a task-based model, Acemoglu estimates that AI's ten-year boost to total factor productivity will not exceed 0.71%, and falls below 0.55% once hard-to-learn tasks are accounted for; he also predicts that the gap between capital income and labor income will widen【verified】. Between micro cost collapse and macro productivity lie organizations, institutions, and a great many tasks that are hard to learn.

05IV. Deeper Implications: Four "Definition-Level" Shifts

**One: the production function must be rewritten.** It moves from Y = F(K, L) toward "reproducible cognitive capital + human judgment + rival physical capital + data + energy"【inference】. Technology is no longer an unbuyable residual but an input that can be produced, owned, and depreciated.

**Two: "marginal cost" is no longer a number but a task cost curve.** It equals capability tier × tokens × steps × unit price, plus verification and trust costs【inference】. Nominal unit price and total task cost can move in completely opposite directions. "Zero marginal cost" is merely a special case of low-difficulty, single-step, already-verified tasks—not a new general rule.

**Three: scarcity has migrated.** When generation becomes cheap, the real constraint becomes "what is worth doing, whether it is done correctly, who is accountable, and whether it can be trusted"【inference】. Scarcity is no longer bits but credibility, judgment, and attention.

**Four: the boundaries of the firm are doubly constrained.** On one hand, falling transaction costs make cognitive outsourcing more worthwhile, so firms can be smaller; on the other hand, rising fixed compute costs concentrate frontier capability heavily【inference】. Organizational form therefore becomes **dumbbell-shaped**: a super-platform plus extremely small trusted units, with the middle layer compressed. The boundaries of the firm are no longer determined solely by the Coasean "internal organizing cost vs. market transaction cost," but jointly by "the externally unbuyable fixed cost of compute vs. the internally non-outsourceable verification and accountability."

06V. Four Action Recommendations for Practitioners

**1. Stop building business models on "unit price."** Build your **task cost curve**: factor in capability tier, number of steps, retry rate, and verification overhead. Any unit-economics model calculated solely on the token unit price will be distorted under real agent workloads【inference】.

**2. Treat verification capability as a core asset, not a cost item.** Since generation is cheap and verification is expensive, "automatable evaluation, regression, auditing, and accountability tracing" is the moat. Whoever turns verification into a reusable asset first will hold pricing power in the age of cost disease【inference】.

**3. Bet on the "non-reproducible" end.** Reproducible cognitive capital will be flattened by full competition; what is genuinely scarce is the specification of ambiguous intent, orchestration across capabilities, backing of outcomes, and taste【inference】. Pile your team's capabilities onto these four things.

**4. Give the macro level time; do not make long-term promises at micro speed.** Cost collapse is a micro fact; the productivity revolution is an unverified macro proposition【verified】. The fundraising narrative can be aggressive; cash-flow forecasts must be conservative.

07Conclusion

The "zero marginal cost" of digital goods is not a revolution—it was already a special case of the old paradigm. The real revolution is four definition-level displacements: **capability becomes a reproducible, depreciable, ownable factor of production; marginal cost is endogenously determined by capability tier, and total cost can rise with efficiency; scarcity migrates from objects to judgment, trust, and accountability; and the boundaries of the firm are doubly constrained by fixed compute cost and the cost of verification and trust.**

The old world's competitive advantage came from "I can produce and you cannot." The new world's competitive advantage comes from "I dare to be accountable for the outcome, while you can only generate."

**Production capability is becoming free. Judgment is becoming the only scarce good.**

This is a living public record. Material revisions will be dated and explained.

Join the inquiry

Add your experience to the discussion

Write a response or simply speak. Peter reviews each contribution before it appears publicly.

DiscussingMarginal Cost Is Breaking Down: When "Cognitive Capability" Becomes a Reproducible, Ownable Factor of Production

Published discussion

0