The Efficiency Mirage: Anthropic and OpenAI Cost Claims Under Scrutiny
Anthropic is raising another $2 billion at a $60 billion valuation. OpenAI follows suit with a $40 billion round. The pitch: their models are cost-efficient enough to charge more than Chinese rivals while still winning on unit economics. The problem? The data trail is cold. No one has seen the receipts.
I have spent the last decade dissecting crypto narratives. From Tezos’s governance collapse to Curve’s veCRON vote manipulation, I learned one law: silence between lines reveals the rot. The current AI cost-efficiency narrative, pushed by a crypto-native outlet like Crypto Briefing, is a structural replay of the same pattern. The claim—that Anthropic and OpenAI achieve superior cost efficiency despite higher prices—is not backed by a single verifiable metric. No training FLOPs, no inference cost per token, no chip-to-chip comparison. Just a punchline designed to land on investor desks.
Let’s dissect the three ways “cost efficiency” can be measured. First, training cost dominance: DeepSeek-V3 trained at 14.8 trillion tokens for under $6 million, while GPT-4 likely cost over $100 million. Second, inference cost per token: OpenAI’s GPT-4o mini runs at $0.15 per million input tokens, while DeepSeek-V2 charges $0.27. Third, total cost of ownership: includes development, energy, and compliance. The Crypto Briefing article does not specify which metric it uses. That is not an oversight. It is a deliberate fog.
My own audit of the Tezos protocol in 2017 taught me to follow the incentives. The project raised $232 million, but I found that the on-chain governance could be bypassed by founders. They called it paranoia. Eighteen months later, $100 million evaporated. Code does not lie, but incentives do. In the AI cost-efficiency debate, the incentive is clear: justify the massive valuation of U.S. AI companies to a crypto audience that wants to believe in “AI as the next DeFi.” The narrative is a weapon, not a report.
What is the real driver of cost efficiency? Chip supply asymmetry. U.S. firms have access to the latest NVIDIA H100/B200 clusters at scale. Chinese firms are restricted to lower-tier A800 or domestic chips. Even if Chinese algorithms are more FLOP-efficient, the hardware advantage can offset that. The article ignores this structural factor. It attributes efficiency to pure algorithm superiority, which is a politically convenient simplification.
I replicated this mental exercise during the 2020 Curve governance scandal. I found that 15% of LPs were being diluted by undisclosed front-running strategies. The industry claimed transparency. The data showed otherwise. Today, the same pattern: claims of efficiency without transparency. The silence between lines reveals the rot.
Now, the contrarian angle: bulls might be right that U.S. models maintain a marginal efficiency edge in English-language inference, especially for complex tasks. But the gap is narrowing. DeepSeek’s Mixture-of-Experts technology and Qwen’s efficient inference stack are closing the distance. If the U.S. edge is 1.5x, the Chinese models offer 3x lower price. For most enterprise use cases, the TCO tilt favors the Chinese side. The real question is not whether U.S. models are more efficient, but whether the efficiency gap is wide enough to justify the price premium.
I do not trust the promise, I audit the perimeter. The perimeter here is the data. Without a public, third-party benchmark that compares cost efficiency across models at the same inference quality level, this entire narrative is a high-priced inventory of unsupported claims. The majority is often the most exploited variable.
My takeaway is simple: if you are allocating capital based on this cost-efficiency story, demand the underlying data. Which benchmark? What timestamp? What inference hardware? What batch size? If the answer is “we can’t disclose,” then you are not investing. You are gambling on a narrative. Governance is not a vote; it is a weapon. And the weapon is currently aimed at your portfolio.