In late 2022, generating a million tokens from a GPT-4-class model cost roughly $20. Today it costs about $0.40 — a 1,000× collapse in just over three years, one of the fastest unit-cost declines in computing history. And yet Nvidia's CEO now stands on stage at GTC and proposes, with a straight face, that the token is the new unit of digital economic value. That revenue equals tokens per watt times available gigawatts.
Both things are true. That's the problem. If your measuring stick shrinks by three orders of magnitude every three years, what exactly are you measuring?
The contenders, briefly
Let's be fair to the candidates. Intelligence — or at least machine utility — has had several proposed units, and each had its moment.
FLOPs dominated the scaling era. Count the compute, predict the capability. Kaplan's scaling laws made this almost respectable science. But FLOPs measure effort, not output. I've seen this exact confusion in PLM programs — teams reporting engineering hours as if hours were features. They aren't. A model that burns 10^26 FLOPs and hallucinates has produced heat, not intelligence.
The watt came next, and it has honest physics behind it. Power is the binding constraint in every data center project I've touched in the last four years — cooling capacity, grid interconnect queues, transformers on 18-month lead times. Huang's "tokens per watt" framing is seductive because it collapses everything into thermodynamic efficiency. But watts measure what you spend, not what you get.
The KV cache — memory footprint per concurrent context — is the engineer's metric, and a good one for capacity planning. Nobody outside an inference team will ever care.
So we arrive at the token. By elimination as much as by merit.
Every model redraws the ruler
Here's the awkward part, and it's where tokenomics gets genuinely slippery. The Stanford AI Index documented a 280-fold drop in the price of equivalent-quality inference — GPT-3.5-tier, MMLU 64.8% — between late 2022 and late 2024 alone. Equivalent quality. Which means the correct unit was never "a token." It was "a token at a fixed quality threshold."
And that threshold won't sit still. GPT-5 and Gemini 2.5 Pro now price input at $1.25 per million tokens — cheaper than GPT-4 was at launch, and categorically smarter. DeepSeek V3 serves at $0.27 per million input tokens. Each new model generation pushes the frontier outward while simultaneously deflating the price of everything behind it. The token of January 2024 and the token of August 2026 are as different as the franc and the euro. Same name, different purchasing power.
I ran into this doing a TCO model for a client last quarter. We priced a document-extraction pipeline at frontier rates, then re-ran it on a distilled open model on their own H100s — $0.18 per million output tokens with vLLM at batch 8, versus a 50× spread against the rented API. Same tokens. Wildly different economics. The word "token" was doing all the work and explaining none of it.
Is there a barometer token, then?
If we must have one, the honest construction is something like: the marginal cost of one token at a fixed benchmark score, on a reference workload, at scale. Epoch AI's more careful work suggests the price for a given benchmark performance falls 5–10× per year — fast, but noticeably slower than the headline 1,000× figures, because the benchmarks keep moving. That gap between the naive number and the quality-adjusted number is precisely where the interesting economics live.
Frontier reasoning still commands $3–15 per million output tokens against $0.03–0.10 for distilled tiers. A 100× spread within the same nominal unit. Any barometer that ignores that spread is marketing, not metrology.wavect+1
The forward look
Watch what happens when agentic workloads make token volume the reported KPI. If tokens per watt per gigawatt becomes how hyperscalers justify $100B campuses — with claimed revenue potential of $150B a year per gigawatt-scale cluster — then we'll have built an entire capital cycle on a unit whose value declines 50% every two months. We've seen this film before. It was called bandwidth, in 1999.yanoai+1
The token may well become the currency of the AI economy. Currencies, though, need central banks, exchange rates, and inflation indices. Who's building the CPI of intelligence? Because until someone does, "we generated 40 trillion tokens this quarter" tells you exactly as much as "we shipped a lot of stuff."
One suspects the answer is: the benchmark that can't be gamed by the next model release. I'm not holding my breath. Are you?
Share: