Edited by humans. Written by AI. How our editing works
All articles

Smarter AI Models May Push Compute Prices Far Higher

Dwarkesh Patel argues that AI revenue growth is outpacing compute supply—and something has to give. That something is probably the price of compute.

Bob Reynolds

Written by AI. Bob Reynolds

August 4, 20268 min read
Share:
A bearded man in a beige shirt gestures while speaking in a home office, with text overlay reading "The end of cheap compute?

Photo: AI. Otieno Okello

Dwarkesh Patel has built a following by asking the uncomfortable questions that the AI industry's promotional apparatus prefers to leave unasked. His latest piece—released as both an essay and a video—is a clean piece of structural reasoning about compute economics, and it's worth working through carefully, because the implications run well past the usual GPU shortage conversation.

The setup is simple. Anthropic's revenue has increased roughly tenfold year-over-year for three consecutive years. Meanwhile, AI lab compute capacity grows at roughly 3x annually. Those two numbers cannot coexist indefinitely without something bending. Patel's question is: what bends?

He identifies three candidates. Lab margins could rise. Compute prices could rise. Or the share of compute devoted to inference—serving paying customers—rather than training could increase. His assessment is that all three are already happening, and that the third option is one the labs actively want to avoid.

That last point is worth sitting with. From the labs' perspective, pivoting compute toward inference is an admission. "If you're spending most of your compute on inference," Patel argues, "you're basically declaring that AI progress has stalled and you're just now in the business of being a cloud provider." That's a less interesting story to tell investors, and a less interesting business to run. The labs believe they're months away from models that make current ones look primitive—and they need to reserve compute for the training runs that get them there.

Which eliminates the third option as a durable solution, and leaves two.

The Margin Ceiling

Could AI labs simply charge more and pocket the difference? Patel is skeptical, and the skepticism is well-grounded. Inference margins for leading models have reportedly climbed sharply—an impressive compression of cost relative to price. But Patel finds it genuinely hard to believe those margins keep rising without attracting competition that competes them back down. "It's just really wild for me to consider that the margins for something like intelligence will be greater than 90% and they don't get competed away at that level."

This is the structural problem with monopoly rents on a product where the underlying capability is improving fast and the number of well-funded competitors is not small. Every quarter that passes gives a new challenger more training data, better hardware, and a clearer target to aim at. Sustainable 90%-plus margins require a moat that is either legally enforced or technically unassailable. Neither condition obviously holds for AI models right now.

So Patel concludes that the real escape valve is compute price appreciation—and he's not treating that as a distant theoretical outcome. GPU spot prices have risen sharply since earlier this year, and the compute supply crunch shows no sign of resolving quickly.

The Google-Anthropic deal with SpaceX illustrates the dynamic at the premium end. Patel notes that Google is reportedly paying $900 million a month for 110,000 GPUs—a blend of GB200s and GB300s—at roughly double the prevailing spot price. When the largest, best-capitalized technology company on earth is paying a 2x premium for guaranteed access to the compute it needs, the spot market is functioning more like a clearinghouse for leftovers than a reliable price signal.

The Alchian-Allen Effect, Applied to Silicon

Here's where Patel's argument gets genuinely interesting. When a scarce, expensive resource is being competed for, the buyers who can extract the most value from each unit outbid everyone else. In compute terms: as GPU time becomes expensive, only the applications that can monetize that time most effectively will be able to afford it.

Patel calls this the Alchian-Allen effect. The economists Armen Alchian and William Allen observed that when a fixed cost is added to two goods of different quality, consumption shifts toward the higher-quality good—because the premium becomes relatively smaller. The same logic applies here. If an H100 costs $20 an hour to rent, you cannot afford to run an inefficient model that burns twice the tokens to produce the same result. You will pay a premium for the model that economizes the compute.

"If you can train the best most efficient model," Patel argues, "then you'll be able to charge much higher margins than you can today." This is the competitive flywheel running in reverse for everyone who isn't at the frontier: as compute gets expensive, the value of efficiency rises, and the labs best positioned to capture that value are the ones that can afford the most compute to train the most efficient models. The rich get richer, and they do it through superior hardware access.

The downstream casualty Patel identifies is the current consumer AI market. "A lot of current popular applications of AI will probably get priced out," he says. Right now, AI is cheap partly because it can't do the most valuable things. As that changes, the willingness-to-pay calculus shifts. An AI that can autonomously conduct research, write production code, or manage complex operations is worth a great deal more per token than one generating marketing copy. The organizations with the highest-value use cases will outbid the lower-value ones. The era of subsidized AI access has a clock on it.

The Lump-of-Compute Fallacy

Patel anticipates the obvious counterargument: if AI creates millions of virtual software engineers, won't the value of any single one of them collapse? He reaches for the lump-of-labor fallacy—the long-debunked idea that there is a fixed amount of work to be done, and more workers only dilute the pool.

Mainstream labor economics rejects that view. The historical pattern, from industrialization through the computing revolution, is that productivity-expanding technologies tend to increase total economic output enough to maintain or raise real wages even as they displace specific jobs. More workers, more specialization, more innovation, more demand. The pie grows.

Whether that pattern holds when the new "workers" are AI agents running at software speed is genuinely unknown. Patel acknowledges as much: "Maybe this labor supply shock will be so big and so fast that we can't count on this general heuristic anymore." The honest answer is that nobody has run this experiment before. What we do know is that every previous wave of labor-saving technology—mechanized agriculture, factory automation, enterprise software—generated more economic activity than it eliminated, eventually. "Eventually" did a lot of work in those transitions, and it may do a lot of work here too.

The Hard Ceiling on Supply

The other side of the equation is supply, and Patel's breakdown of why 3x annual compute growth is both hard to maintain and nearly impossible to accelerate is among the most concrete pieces of the argument.

That 3x comes from three sources. Moore's Law accounts for roughly 1.4x—and Patel thinks sustaining even that will be difficult. New semiconductor fabs contribute roughly 1.2x, but that pipeline is gated by ASML's production of extreme ultraviolet lithography machines, a constraint that won't resolve before 2030. The remaining 1.8x comes from AI absorbing wafer capacity previously allocated to smartphones and PCs—a finite reallocation that approaches its ceiling as AI's share of leading-edge node production approaches saturation.

Add those up and you don't obviously get to 3x for much longer, let alone beyond it. The capital expenditure commitments that Google, Microsoft, and Amazon have made suggest they understand this—you don't sign nine-figure monthly compute contracts unless you believe the spot market will be worse, not better.

Patel closes with an honest self-critique. He notes his analysis rhymes with historical resource-scarcity arguments that turned out to be wrong—Paul Ehrlich's famous bet against Julian Simon on commodity prices being the canonical example. Ehrlich lost. But Patel's rebuttal is structural: the supply of compute is far less elastic than the supply of metals. You cannot find a substitute for ASML EUV machines the way you can switch from copper wire to fiber optic cable. The physical and industrial constraints are specific and documented.

That honesty about the limits of his own argument is what makes it worth engaging with. Patel isn't predicting compute will cost ten times more. He's showing that the arithmetic of current trends points hard in that direction, identifying the mechanisms that could produce that outcome, and noting—correctly—that the supply-side is not well-positioned to absorb the shock.

He ends with a discomfort that feels genuine: "I wish we didn't live in a world with such strong economies of scale for intelligence because I'm worried about power concentration, but it seems we do."

That's not a prediction. It's a problem statement. And it's one that deserves more attention than the industry's quarterly earnings calls tend to provide.


Bob Reynolds is Senior Technology Correspondent at Buzzrag.

From the BuzzRAG Team

AI Moves Fast. We Keep You Current.

Framework breakdowns, tool comparisons, and AI coding insights — distilled from the best tech YouTube creators. Free, weekly.

Weekly digestNo spamUnsubscribe anytime

More Like This

Smiling man in green shirt points to a window displaying the /routines app logo with API, webhook, and schedule options

Anthropic's Claude Routines Targets No-Code Automation Market

Claude Routines lets users automate workflows with natural language instead of drag-and-drop builders. Is this the end of traditional no-code platforms?

Bob Reynolds·4 months ago·6 min read
Metallic robotic figures with glowing spherical heads against a dark background, with "SUB-AGENTS" text overlaid in white

AgentZero's Sub-Agents: Self-Modifying AI Delegation

AgentZero demonstrates AI agents that create and manage specialized subordinates on demand. The system modifies itself—which raises practical questions.

Bob Reynolds·5 months ago·6 min read
Google Cloud logo with two smiling engineers holding a device in a lab setting, text reads "Should I even use AI?

Not Every Problem Needs AI. Here's How to Tell.

Google engineers explain when to use generative AI, traditional machine learning, or just plain code. The answer matters more than you'd think.

Bob Reynolds·5 months ago·6 min read
Comparison showing Opus app costing $100 crossed out, with arrow pointing to Advisor app costing $1, featuring circular…

Anthropic's Advisor Strategy: When Cheaper AI Models Work Better

Anthropic's new advisor strategy pairs expensive Opus with budget models, cutting costs by 12% while maintaining quality. But testing reveals surprises.

Bob Reynolds·4 months ago·5 min read
Colorful tech logos (Linux penguin, whale, palm tree icons) with neon gradient background and smiling woman in baseball cap…

DeepSeek V4 Undercuts AI Giants While France Ditches Windows

DeepSeek's V4 slashes AI inference costs by 90% as France commits to Linux migration. Plus: Ubuntu's local inference push and Linux drops 486 support.

Yuki Okonkwo·3 months ago·6 min read
Bearded man wearing glasses and white beanie adjusts his frames against dark background with bold text reading "THEY MISSED…

AI's Inference Crisis: Why Sora Died Burning $15M Daily

OpenAI killed Sora after six months. The reason reveals AI's shift from training races to inference economics—and what breaks next.

Marcus Chen-Ramirez·4 months ago·7 min read
Two instructors in Jedi robes holding lightsabers stand before a Cisco CCNA 200-301 course interface with network diagrams…

NetworkChuck's Free CCNA Program Draws 35,000

NetworkChuck and Jeremy Ciorra launched a free CCNA program that drew 35,000 signups. Here's what the model actually offers—and what it reveals about online learning.

Bob Reynolds·3 months ago·7 min read
Apple devices and AR glasses displayed against a colorful gradient background with text reading "Apple's next big thing

Apple's Ultra Strategy: Premium Tier or Price Ceiling?

Apple plans to expand its Ultra lineup beyond watches to iPhones, MacBooks, and AirPods. What this means for pricing and innovation across product tiers.

Bob Reynolds·3 months ago·5 min read

RAG·vector embedding

2026-08-04
1,882 tokens1536-dimmodel text-embedding-3-small

This article is indexed as a 1536-dimensional vector for semantic retrieval. Crawlers that parse structured data can use the embedded payload below.