U.S. firms shift to cheaper Chinese AI models

Rising 2026 usage bills pushed U.S. companies to cheaper Chinese models; Uber exhausted its AI budget in four months and Chinese models now exceed 30% of OpenRouter token use.

U.S. companies are increasingly routing tasks to lower-cost Chinese AI models after consumption-based billing drove sharp cost increases in 2026. One major ride-hailing firm exhausted its annual AI allocation within four months after widespread adoption of coding tools, prompting managers to limit usage. On OpenRouter, weekly token use of Chinese models has stayed above 30% since early February and has reached peaks near 46% this year.

The shift reflects a search for similar capability at lower prices as firms deploy AI for software development, customer service and automation. While per-token prices have fallen in many cases, overall bills rose when vendors moved from flat subscriptions to pay-as-you-go models. Teams now tend to send routine or less demanding work to cheaper models and reserve more expensive models for complex tasks.

Price gaps are large. Some Chinese models charge the equivalent of a few dozen cents per million tokens compared with roughly $4 per million for leading U.S. frontier models. One provider, DeepSeek, charges about 3% of the token price of a leading U.S. model. Zhipu AI’s GLM 5.2, launched in June, has been adopted quickly; when hosted in China its token cost runs around 15% of comparable U.S. models.

Developers and startups point to cost as the main driver. Harpreet Arora, head of agentic infrastructure at Vercel, described pricing as “doing the work here,” noting that teams route tasks to the cheapest model that is good enough. San Francisco startup Lindy reported switching from pricier models to DeepSeek and said the change saved the company millions. One developer said they use expensive U.S. models for complex planning while running routine coding and voice recognition on lower-cost Chinese models.

Major cloud providers have made Chinese models easier to access by offering them through their platforms. Competition among Chinese developers has intensified, with firms including Alibaba, Moonshot AI and Zhipu racing to improve performance and benchmark rankings. GLM 5.2 has been among the fastest-growing models in daily token volume and customer count in recent tracking.

Security and regulatory concerns remain. One prominent Chinese provider was added to a U.S. trade blacklist in 2025. Regulated industries and companies handling sensitive data are cautious about sending information through systems hosted abroad. GLM’s open-weight architecture can be deployed on a firm’s own servers or private cloud, allowing model use without sharing data with the developer.

Analysts and companies say the current price gap may narrow. Some U.S. providers are considering significant price cuts amid growing competition, and enterprises are weighing cost, capability and risk when building AI stacks. A major Chinese developer has been reported to be considering a large share sale after a sharp post-listing rally; a six-month IPO lock-up is due to expire in early July.

Val Bercovici, chief AI officer at WEKA, described the trade-off in direct terms, saying open-source models can be “90% as good at 10% of the price,” which makes lower-cost options the preferred choice for much everyday work while premium models are used for the most demanding tasks.

Articles by this author