An interactive explainer · ~6 min read

Why everyone's suddenly talking about an AI bubble.

Every $20 chat, every $200 power-user plan, every trillion-dollar data-center deal is wired to the same thing: a token costs real compute, and most of that compute is being sold to you for less than it costs to make. This is what people mean when they say the subsidy era is ending.

2025 OpenAI revenue $13.1B Tripled year over year
2025 OpenAI cash burn $9B A 70% burn rate
Adjusted gross margin 33% Healthy SaaS runs ~75–80%
Projected spend, 2025–2030 $665B $440B of it on training

Sources: OpenAI 2025 financials · Anthropic at $380B · Bloomberg circular deals

01 · The Loop

The money goes in a circle. That's the first reason people are nervous.

Nvidia writes a check to OpenAI. OpenAI writes a check to Oracle. Oracle writes a check back to Nvidia. The same dollar shows up as investment, revenue, and capex on three different income statements. Stocks rally on each leg.

The AI capital loop A circular diagram showing money flowing from Nvidia to OpenAI as investment, from OpenAI to Oracle as cloud spend, and from Oracle to Nvidia as chip purchases. Nvidia CHIPMAKER OpenAI AI LAB Oracle CLOUD INVESTS up to $100B CLOUD SPEND ~$300B over 5 yrs BUYS CHIPS $40B+ orders SAME DOLLAR · COUNTED 3×
Capital flow

Bulls call this a virtuous circle — chipmakers locking in customers, AI labs locking in supply. Bears call it vendor financing dressed up as growth, the same pattern that powered the 1999 telecom boom right before the bust. Both are technically correct. The disagreement is about whether the end demand — you, paying for ChatGPT or Claude or Cursor — actually justifies the buildout.

A token is just compute, billed by the syllable.
Chapter 02

02 · The Unit Economics

What a token costs — and why your $20 plan is the variable.

Models charge by the token (roughly ¾ of a word). Input is cheap. Output is 3–5× more expensive because each output token requires a fresh forward pass through tens of billions of weights. These are the real-world wholesale prices labs charge developers.

Claude Haiku 4.5 SMALL / CHEAP
$1 / $5
Per million in / out tokens

Used for routing, classification, fast chat. About 0.0001¢ to read a short message.

Claude Sonnet 4.6 WORKHORSE
$3 / $15
Per million in / out tokens

The default for serious work — coding assistants, research, agents.

Claude Opus 4.8 FRONTIER
$5 / $25
Per million in / out tokens

Top-of-stack. A single agent run that reads a codebase can easily burn $0.50–$5 of Opus.

Opus 4.8 · Fast Mode PREMIUM
$10 / $50
Per million in / out tokens

2× the price for speed. This is what your subscription is reaching for when you click "think harder."

Anthropic's public pricing is the closest thing to a wholesale market for intelligence. Now hold those numbers in your head and look at what consumer plans actually charge.

03 · The $20 Paradox

Drag the sliders. See where the subsidy starts.

A typical Pro subscription is $20/month. At wholesale token prices, here's what a heavy user actually costs to serve. Numbers update as you change them.

Casual chat ≈ 5/day · power user ≈ 100/day · agent loops ≈ 300+/day
Short Q ≈ 500 · coding session ≈ 8,000 · agent w/ tools ≈ 40,000+
A paragraph ≈ 200 · a long answer ≈ 1,500 · code generation ≈ 3,000
Tokens per day
Tokens per month (30d)
True API cost (wholesale)
What you pay $20
Lab's gross margin on you
Verdict
You're profitable for them.

These are wholesale API prices — the rate Anthropic or OpenAI would charge a developer to do the same work. The labs' actual cost-of-goods is lower (they own the compute), but it's not zero. SemiAnalysis estimated ChatGPT at ~0.36¢ per query back when models were smaller; today's frontier models cost meaningfully more per call. The shape of the loss is the same.

Your $20 plan is a bet that you won't use it.
Chapter 04

04 · The $200 Paradox

So why wouldn't you just pay $200 for Perplexity instead of Claude?

On the surface, two $200/month plans look identical. Both promise frontier models. Both promise "unlimited." Under the hood they're different bets with different exit hatches.

Claude Max 20× $200 / mo
  • Direct access to Anthropic's own models, no router in between.
  • ~900 prompts per 5-hour window, deep Opus & Sonnet quotas.
  • Claude Code CLI — agentic coding sessions that can burn 30,000+ tokens per turn.
  • Weekly compute-hour caps (≈40 Opus hrs, ≈480 Sonnet hrs).
  • One vendor. One model family. No GPT-5 or Gemini.
  • Limits tightened twice in 2025–26 — the subsidy is visibly contracting.

You're paying for raw token throughput against one lab's frontier models.

Perplexity Max $200 / mo
  • Routes across 19+ models — Opus, GPT-5, Gemini, Grok, in one interface.
  • Unlimited Pro search + unlimited Labs + Comet browser + 10K Computer credits.
  • Heavy caching: identical web pages get fetched once and reused across users.
  • Frontier-model usage is generous but not literally unlimited.
  • A router decides which model answers — you don't always get Opus.
  • Perplexity is paying API prices for every uncached call you make.

You're paying for orchestration and aggregation across many labs' models.

The honest answer to "why not just pay Perplexity?" is: for most people, you should.

If you don't need a CLI agent burning tokens for hours on your repo, Perplexity Max gives you broader model access and more product surface for the same money. If you live inside an agentic coding loop, Claude Max gives you a deeper token bucket against one specific frontier. Both companies are losing money on heavy users — Perplexity by paying upstream API bills, Anthropic by serving compute below cost. They're just losing it differently.

05 · The Burn

And this is what "subsidy era ending" looks like at the company level.

OpenAI's own projections through 2030. Revenue grows fast. Cash burn grows faster. For most of this window, every additional dollar of revenue is paired with a dollar (or more) of fresh cash being set on fire.

Source: OpenAI 2025 actuals and internal 2026–2030 projections, as reported via leaked financials in March 2026. Anthropic projects breakeven by 2028 with burn dropping to 9% of revenue — a much tighter path.

The bull case: inference costs drop ~10× per year as chips and algorithms improve. If usage grows slower than costs fall, margins go from 33% to 75% without anyone raising a price. The bear case: usage is growing faster than costs fall, frontier-model training cost is going up, and the entire structure is propped up by the circular financing in Chapter 01.

06 · The Verdict

Is it actually a bubble?

The honest answer is: both things are true at once. The technology is real and the revenue is real. The valuations and the financing structure may not be.

Why "bubble" rings true

The financial plumbing looks fragile.

  • Vendor financing: Nvidia investing in customers who use the cash to buy Nvidia chips. The pattern that preceded 2000.
  • Anthropic is valued at 27× revenue at 40% gross margins. Pure software at the same multiple runs 75%+ margins.
  • Consumer plans are demonstrably unprofitable on power users — Cursor proved it publicly in mid-2025 when it had to reprice.
  • OpenAI's own projection: $665B of spend through 2030, $440B on training alone. That requires near-perfect execution and continued capital access.
  • One frontier-model release that doesn't deliver, or one round that fails to close, and the loop breaks.

Why "bubble" misses something

The underlying usage is not fake.

  • Anthropic went from $1B to $14B in annualized revenue in 14 months. Real customers, real contracts, mostly enterprise.
  • Inference cost-per-query has fallen roughly an order of magnitude per year for three years running.
  • Models keep getting better at things that translate directly into labor savings — coding, research, support, ops.
  • Even if half the AI startups die, the survivors plus the hyperscalers absorb the demand and the chips.
  • The 2000 dot-com bust still gave us Amazon, Google, and the modern internet. Bubbles can be real and productive.

What "the subsidy era is ending" actually means.

It doesn't mean ChatGPT vanishes or Claude raises prices to $500/month overnight. It means a slow, visible re-pricing: usage caps tightening, power users pushed to usage-based billing, smaller models doing more of the work behind a flagship-model wrapper, and "unlimited" quietly becoming "fair use."

Cursor's pricing change in July 2025 was the canary. Anthropic's 5-hour and weekly limits tightening through 2025–26 is the next data point. Watch what your favorite tool's pricing page looks like a year from now — that's the real chart.