Token Costs Must Drop 90% for Scaling: The Hidden Cliff of Technology Adoption
When OpenAI boasts of "54% token efficiency gains" but Palo Alto Networks' CEO coldly responds "that's not nearly enough—costs must fall 90%"—why can't incremental progress reshape markets, and why does only a cliff-like cost collapse truly enable scale?
8 min read
The Event
In July 2026, Palo Alto Networks CEO Nikesh Arora made a striking assertion on CNBC: even though OpenAI's latest GPT-5.6 "Sol" model achieved 54% improvement in reasoning efficiency, enterprise AI deployment at scale still requires a 90% reduction in token costs over the next two years to be considered viable. He expected prices to decline by up to 20% in the first year, followed by an "explosive" expansion of cost declines reaching 90% thereafter.
OpenAI CEO Sam Altman views this as a technological milestone. But Arora's response was cold: it's "only a good start." The implicit message is clear—in enterprise procurement logic, incremental improvement simply doesn't cut it.
Why 54% Progress "Isn't Enough"
On the surface, a 54% efficiency gain looks like technological progress. But enterprise finance departments do the math simply:
- An organization cutting AI inference costs by only 54% must rely on "using it smarter" to balance spending, meaning the marginal revenue per token generation must more than double.
- In other words: efficiency gains are cancelled out by expectations. Companies have become accustomed to 20-30% new capabilities every six months; this efficiency improvement is actually reinterpreted as "decline"—because it fails to fundamentally shift costs.
- Once costs fail to reach the "scalable" threshold in enterprise financial models, decision-makers postpone deployment. Delayed purchases = vendors don't get paid = infrastructure investments can't be recovered.
The Economics of Critical Mass
Geoffrey Moore described a similar phenomenon in *Crossing the Chasm*: the inflection point between Early Adopters and the Mainstream Market isn't "slightly better products"—it's reaching a qualitative threshold where user experience shifts from "only professionals can handle this" to "anyone can use this."
The AI cost threshold follows similar logic:
1. Early Stage (2022-2024): Enterprises gladly pay a 50x cost premium just to access GPT-4, because the marginal benefit of capability is enormous (the leap from "nothing" to "something").
2. Transition Stage (now): Capability gradually becomes "commoditized" (OpenAI, Claude, Gemini, DeepSeek are all converging). Competition shifts from "can we do it" to "what does it cost." A 54% efficiency gain can't change this landscape—because equally capable competitors are also improving.
3. Scale Stage: Only when costs drop enough to move enterprise financial models from "difficult ROI" to "obvious ROI" does mass deployment explode. This inflection point is marked by a single budget that previously bought 100 tokens now buying 1,000 or even 10,000 tokens.
A 90% cost reduction means shrinking costs to 1/10th—that's not improvement, that's rewriting the rules.
Enterprise Procurement's Hidden Decision Model
Palo Alto Networks is a cybersecurity software company with $8 billion in annual revenue and direct contact with thousands of large enterprises. Arora knows what these companies are actually asking—not "how smart is GPT-5.6?" but:
- Our annual IT budget is X; if each AI call averages $0.001, how much will this cost annually?
- If costs don't drop, how many internal applications can we run inference on simultaneously?
- If we switch from OpenAI to cheaper open-source models, how much does quality degrade? (= Decision-makers start making price-quality trade-offs)
Once enterprises begin trading off "OpenAI vs. open source," frontier models lose. Because open source is free—only computation costs apply.
The true meaning of 54% efficiency gains: Computation costs aren't dropping enough, and enterprises are reconsidering open-source alternatives. That's a danger signal.
The Industry Chain Hostage Situation
Sequoia Capital estimates that global AI infrastructure spending reaches massive scale by 2026. But these investments' IRR (internal rate of return) is bottlenecked by "token prices can't drop enough."
- Nvidia, hyperscale cloud providers (AWS, Azure, Google Cloud), and model companies like OpenAI all need scale to amortize costs.
- But scale requires demand. Demand requires enterprises to see costs as "worthwhile."
- If token costs only drop 20%, enterprises still see it as "not worthwhile," delaying purchasing decisions.
- Delayed purchasing = infrastructure investments can't be recovered = the industry chain's cash flow breaks.
In other words, Nikesh Arora isn't boasting—he's warning: the entire AI industry now sits on the cliff edge of "over-investment, under-returns." Only a 90% cost drop can trigger "genuine" scale, can turn all investments from losses to gains.
Analogies and Lessons
This pattern repeats throughout technology history:
- Solar panels: From the 1980s through 2000s, efficiency improved 20-30% per decade, but markets remained "niche applications" (satellites, remote charging). After 2010, costs fell 80%, and the market exploded into a global energy industry. Those 20 years of "efficiency improvement" actually didn't matter.
- Electric vehicles: When the Model S launched (2012), it was a billionaire's toy with terrifyingly expensive batteries. Battery costs declined steadily at 7-10% annually; by 2020, battery costs had fallen enough to make the Model 3 "an everyday car." But the inflection from "billionaire's toy" to "mass market car" wasn't about efficiency—it was about cost critical mass. Once total vehicle price approached gas car parity, the market exploded.
AI stands at the same inflection point now.
The Hidden Risk
Nikesh Arora's warning also reveals a "hidden truth": OpenAI, Google, Anthropic all claim "models are getting smarter," but nobody dares plainly state "but inference costs can't drop significantly." Because once that's said, market pricing logic reverses—from "pay premium for future capability" to "discount for current costs."
If model companies' marginal costs (computation + electricity) already approach 80-90% of sale price, there's no room to cut prices further. Unless they either invest in new technology (inference quantization, distillation, etc.) or accept massive profit margin erosion.
That's why Arora said "only a good start"—he's hinting: frontier model companies are now trapped. Don't cut prices and you can't survive; cut prices and you hollow out profits.
Preparing your check…
Source: TechOrange