The pricing move that landed this week has been framed as a discount war. It is not. When Google pushed the message that enterprise AI costs are too high, it was not running a promotion. It was demonstrating a structural advantage that no lab renting compute can replicate: it owns the silicon the inference runs on.
The numbers tell the story bluntly. Gemini 3.1 Pro undercuts the flagship rates from Anthropic and OpenAI on input, and it pairs that with a 1M-token context window. But the specific “60% cheaper” claim and several of the named comparisons in the original framing do not hold up on current published rate cards. As of late August 2026, Anthropic lists Claude Opus 4.8 at $5 per million input tokens, and OpenAI lists GPT-5.5 at $5 per million input tokens. Google’s published Gemini pricing varies by tier and SKU, including a preview rate that can be lower than those figures, but the exact percentages and “largest context window” statement depend on which products you are comparing and are not consistently true across official pages.
Google can price here because it is not paying Nvidia’s margin on the same terms as everyone else. Vertical integration allows Google to sidestep a layer of third-party accelerator economics, which can translate into a real cost advantage. At Google Cloud Next in April 2026, the company unveiled TPU 8i, its dedicated inference chip. Google’s own technical material says TPU 8i delivers up to 80% better performance per dollar for inference than the prior generation. Gemini is trained on TPUs and is primarily served on TPUs, but it is not accurate to state it is served exclusively on TPUs in all cases.
The labs on the other side of this equation are not standing still. Salesforce and Anthropic announced a broader partnership in late August 2026, but the original draft’s claim that Salesforce committed approximately $300 million on Anthropic tokens in 2026 alone, and that it has an existing $300 million equity stake in Anthropic, is not supported by the partnership announcement or public filings that are easily verifiable. What is supported is the structural point: for a lab that relies heavily on external infrastructure, a meaningful share of revenue flows back out through compute and hosting. In October 2025, Anthropic announced an expansion with Google Cloud that included plans to access up to one million TPUs, described as worth tens of billions of dollars. Anthropic’s biggest partner on infrastructure can also be a direct competitor at the model layer.
Meanwhile, the threat that spooked Western labs for most of 2026 is correcting. DeepSeek’s own API documentation now shows peak and off-peak pricing, including a move in peak-hour output pricing from $0.28 to $1.32 per 1 million output tokens for one of its offerings. The broader point remains intact: the first race to the floor was hard to sustain once usage scaled and the full training and serving bill showed up.
This matters enormously for how investors should think about the pure-play labs. Recent reporting, citing the Financial Times, has described some Anthropic backers modeling an IPO as early as October 2026 at $2 trillion or more, alongside expectations for a $100 billion to $120 billion annualized revenue run rate by the end of 2026. That is extraordinary growth. But a $2 trillion valuation prices in terminal value, and terminal value requires durable gross margins. When your largest competitor owns a major part of the cost curve for tokens, margin durability becomes a genuine question.
Anthropic’s answer is differentiation on capability and trust: the Salesforce partnership is built on Claude’s reasoning depth, not its price. Customers are paying a premium today. But the original claim that Anthropic’s top model is priced more than 2.5 times higher than OpenAI’s flagship is not consistently true on the standard per-token API rates: Anthropic’s Opus 4.8 lists at $5 per million input tokens while OpenAI’s GPT-5.5 lists at $5 per million input tokens. Where the premium can show up is in other parts of the bill, including which model you treat as “top,” output-heavy workloads, and contractual enterprise packaging, which are not captured by a single headline ratio.
For holders of Alphabet, the thesis is straightforward. The silicon moat is not speculative; it is shipping, and Google has published specific price-performance claims for TPU 8i. For anyone evaluating Anthropic ahead of an October listing that investors are discussing publicly, the harder work is stress-testing what the income statement looks like at scale, inside a market where a price setter also owns a big piece of the infrastructure stack.
