Index  ›  ai  ›  Forbes

As Token Costs Plunge, Enterprise AI Providers Face A New Margin Squeeze

Forbes Published Jul 28, 2026 Reviewed Jul 28, 2026 ✓ Reviewed by citations.press editors
As Token Costs Plunge, Enterprise AI Providers Face A New Margin Squeeze
Anthropic earned about 80% of its $30 billion annual run-rate from enterprise use and saw its inference margins rise from 38% to over 70% in 2026, according to SemiAnalysis.
about 30000000000 USD · Anthropicabout 80 % · Anthropic38 % · Anthropic inference marginsat least 70 % · Anthropic inference margins
Chinese AI model suppliers including DeepSeek, Qwen, GLM, Kimi and MiniMax captured 46% of the generative AI model market share by mid-2026.
46 % · Chinese AI model suppliers including DeepSeek, Qwen, GLM, Kimi and MiniMax
According to Morgan Stanley, Amazon, Microsoft, and Google are expected to spend $805 billion on capital expenditures in 2026 and up to $1.1 trillion in 2027.
805000000000 USD · Amazon, Microsoft, and Googleat least 1100000000000 USD · Amazon, Microsoft, and Google
Nvidia reported a net margin of 63% in the period covered by the article.
63 % · Nvidia
According to Cursor research cited by The Wall Street Journal, building a web browser using Cursor’s Composer and Anthropic’s Opus 4.8 costs $1,339, compared to over $10,000 on OpenAI’s GPT-5.5.
more than 10000 USD · OpenAI’s GPT-5.51339 USD · Cursor’s Composer and Anthropic’s Opus 4.8
AWS reported a 37.7% operating margin and Google Cloud reported a 35.6% operating margin.
37.7 % · AWS35.6 % · Google Cloud
TSMC, SanDisk, and Micron reported net margins of 46.5%, 34.2%, and 56%, respectively.
46.5 % · TSMC34.2 % · SanDisk56 % · Micron
OpenAI generated approximately $25 billion in annualized revenue and operated with a gross margin of about 33%.
about 25000000000 USD · OpenAIabout 33 % · OpenAI
Pylon, a customer support platform for high voltage towers, received $1.6 million in free tokens from one vendor, $65,000 from another, and $10,000 from a third, according to co-founder Marty Kausas.
about 1600000 USD · Pylon65000 USD · Pylon10000 USD · Pylon
CoreWeave had deeply negative free cash flow, a roughly $99 billion backlog, and about $24.9 billion of debt.
about 99000000000 USD · CoreWeaveabout 24900000000 USD · CoreWeave

AI spending cuts are driving a rapid collapse in token prices, forcing generative AI model providers into a significant pricing reckoning. Companies are increasingly choosing cheaper, often Chinese, models for routine tasks, avoiding the overuse of expensive, powerful AI. This shift, dubbed "tokenomics," is squeezing profit margins for costly U.S. LLM providers like OpenAI and Anthropic, while boosting Chinese rivals and chip manufacturers such as Nvidia. The industry's profit potential is diminishing for chatbot suppliers due to intense competition, strong supplier power, and empowered buyers. Investors are urged to monitor token prices and inference margins as the market reconfigures, favoring a hybrid model approach for efficiency.

Companies are slashing AI spending, triggering a rapid collapse in token prices that is squeezing margins across the generative‑AI stack and forcing model providers into their first real pricing reckoning — with major implications for investors watching the sector’s profitability evaporate.

The most powerful and expensive AI models aren’t necessary for relatively mundane tasks. “It’s like driving a Lamborghini to go to the grocery store to pick up milk when that was designed to be raced around a track,” Cursor field chief technology officer Mike Saeks told Wall Street Journal" href="https://www.wsj.com/business/china-us-ai-model-costs-53a12e96?mod=hp_lista_pos2" rel="nofollow noopener" target="_blank">The Wall Street Journal.

In response, companies are using lower-priced models, including some built in China. The potential cost savings — for example, saving 87% on the cost of building a web browser — are considerable.

Doing that from scratch costs more than $10,000 on OpenAI’s GPT-5.5; whereas a combination of Cursor’s Composer and Anthropic’s Opus 4.8 gets the job done for $1,339, according to Cursor research featured by the Journal.

Businesses that use such chatbots are getting big discount offers. Pylon, a customer support platform for high voltage towers, has received around $1.6 million in free tokens from one vendor, $65,000 from another and $10,000 from a third, co-founder Marty Kausas told the Journal.

The fierce price competition is bringing to life a new field of study for decision-makers: tokenomics. This discipline helps businesses make the best use of limited AI budgets by tracking the rapidly changing price of tokens — chunks of data on which AI chatbot prices are set — as well as the cost of providing those tokens to users.

By applying tokenomics, AI users are inadvertently creating winners and losers across the generative AI value network.

The winners include Chinese suppliers of less expensive large language models — including DeepSeek, Qwen, GLM, Kimi and MiniMax — which by mid-2026 won 46% market share. Other beneficiaries are GPU and memory chip designers and fabricators such as Nvidia, TSMC, Micron, SanDisk and Western Digital.

Meanwhile, more expensive LLM providers — such as Cursor, Uber and Salesforce — are adjusting their prices downward.

My work with Harvard Business School professor Michael Porter — whose books in industry analysis and competitive advantage transformed the field of business strategy — is making me acutely sensitive to the forces driving down the AI industry’s profit potential.

I previously mapped out this industry — identifying segments that vary in their inherent profitability. In a nutshell, the AI chip design and manufacturing segment is highly profitable while the AI chatbot suppliers lose huge amount of money.

How so? OpenAI and Anthropic face attacks from all sides. Suppliers have huge bargaining power; price competition is fierce as is pressure to innovate; the cost for new entrants to participate is dropping; and customers face low switching costs and a strong incentive to find the best value.

With IPOs looming for OpenAI and Anthropic, the diminishing profit potential of the AI chatbot market does not bode well for retail investors.

The chip and cloud services providers enjoy the highest profit potential while the model providers are being squeezed the most.

Model providers serving consumers suffer lower margins than business-focused companies. OpenAI, at roughly $25 billion of annualized revenue, is estimated to run about a 33% gross margin — dragged down by consumer ChatGPT, which charges a monthly rate regardless of how much is consumed.

Anthropic — which gets about 80% of its $30 billion run-rate from enterprise use — saw inference margins climb from 38% to over 70% during 2026 due to the popularity of Claude Code, according to SemiAnalysis.

AI software costs more than cloud-based software. Providers have responded with price changes. For example, heavy agentic use forced Cursor to replace its $20 flat plan with credits-based repricing. Uber used up its entire 2026 budget in four months. Salesforce has changed its Agentforce pricing models three times in eighteen months.

AI cloud services providers — with the exception of CoreWeave — are profitable. CoreWeave has deeply negative free cash flow, a roughly $99 billion backlog and about $24.9 billion of debt.

Other hyperscalers are part of larger companies — such as Amazon, Microsoft and Google — which are expected to spend $805 billion on capital expenditures in 2026 and as much as $1.1 trillion in 2027, per Morgan Stanley.

AWS is profitable (37.7% operating margin) as is Google Cloud (35.6% operating margin). It is difficult to tell whether Microsoft’s AI cloud services are making money since it bundles Azure and AI infrastructure into its broader Intelligent Cloud segment.

Chips are highly profitable. Nvidia’s net margin is 63%. Other chip makers are also highly profitable — such as TSMC (46.5%), SanDisk (34.2%) and Micron (56%). With demand vastly exceeding supply and new production capacity not coming online for at least a year, prices and margins are rising.

The winners are companies that own the token-cost floor or capture the deflation. Nvidia, TSMC and the high-bandwidth memory companies such as SanDisk benefit as cheaper tokens drive more demand. Google is uniquely benefiting from both custom silicon — helped by a deal with Anthropic to use the company’s Tensor Processing Units — and distribution.

Anthropic benefits from its enterprise focus, and users of AI are getting more bang for their buck by routing routine work to a $2 model while reserving a $25 model for the most difficult 20% of their computing tasks..

The losers suffer from the gap between flat-rate pricing and unhedged token costs. Indebted neoclouds with single-customer concentration are at risk, as are leading U.S. frontier model providers, which are just 2.7% ahead of China’s best, noted Stanford’s 2026 AI Index.

Ultimately the future of the AI industry depends on whether the value of AI exceeds its price. For now, companies are working on lowering the price.

A case in point is Telnyx, whose AI agent costs were $100,000 a day using Anthropic. This cost sent the company to open models. “They worked,” CEO David Casem told the Journal. “It’s not like we don’t use OpenAI or Anthropic models, we still do. They just don’t do everything anymore,” he added.

The company’s 1,400 agents use Chinese startup Z.AI’s product which costs around $100 per agent per day. Anthropic’s most powerful model, Fable, acts as a conductor that plans out work while open-weight models do the implementation. OpenAI’s Sol handles a review of what the open-weight models produce, added the Journal.

Investors should consider placing their bets on the companies with the most profit potential in light of the forces driving Telnyx’s AI consumption recipe.

This article was originally published by Forbes ↗. citations.press indexes the source-backed facts above and links to the original. Something wrong? Corrections policy · Report an error