Why DeepSeek Costs 25x Less: Creative Destruction Replayed in Semiconductors
When OpenAI charges $20/M tokens and a Chinese startup delivers comparable LLM quality at 1/25 the price—why did the market reprice NVIDIA overnight?
5 min read
The Event
DeepSeek released its R1 model in early 2025, achieving performance close to OpenAI o1 across multiple benchmarks, with training costs of only $6M (vs. OpenAI's estimated $80M+) and API inference pricing at $0.55/M output tokens (vs. OpenAI o1's $60/M).
Why This Is Creative Destruction, Not Incremental Improvement
Schumpeter wrote: true innovation is "not making carriages faster, but inventing the train." DeepSeek didn't try to make transformer training faster—they rethought the entire stack:
- MoE (Mixture of Experts) architecture: activate only needed parameters, slashing inference costs
- Multi-head latent attention: compress KV cache, break through memory bottlenecks
- FP8 training: cut hardware costs in half
- Public reasoning traces: make fine-tuning no longer dependent on OpenAI APIs
Each is an architectural shift, not hyperparameter tuning.
Why NVIDIA Dropped 17% Overnight
If a task that previously required 100 H100s now needs only 4—the market's implicit assumption that "AI always needs more GPUs" collapses instantly. This isn't NVIDIA's product getting worse; it's the market repricing "future GPU demand."
But Schumpeter also said: creative destruction doesn't eliminate demand, it reorganizes it. Jensen Huang himself responded: "Cheap inference will unleash 10x more use cases"—that's Schumpeter's logic too.
Preparing your check…
Source: Import AI