Chinese open weight AI models now handle almost a third of all enterprise AI traffic on one major platform

Started by Kev5, Jul 16, 2026, 04:26 AM

Previous topic - Next topic

0 Members and 1 Guest are viewing this topic.

Topic: Chinese open weight AI models now handle almost a third of all enterprise AI traffic on one major platform   Views(Read 118 times)

Kev5

Open weight AI models, many built by Chinese companies, processed 29 percent of all AI tokens routed through cloud platform Vercel's production gateway in June, up sharply from roughly one ninth of total volume back in April, according to Vercel's latest AI Gateway Production Index. Despite handling nearly a third of all traffic, these models accounted for less than 4 percent of total spending, reflecting costs roughly a tenth of the platform average

DeepSeek was the biggest beneficiary of the shift, capturing 22.6 percent of token volume in June, making it the third largest provider on the gateway behind only Anthropic and Google, and closing to within two percentage points of Google's slipping second place position. Vercel's head of agentic infrastructure Harpreet Arora summed up the dynamic simply, when a task doesn't need the best model available, teams are increasingly routing it to the cheapest one that's good enough, and the recent wave of Chinese models is winning that particular trade

Overall AI investment kept climbing through June, with token volume up 29 percent and spending up 27 percent, but unlike the prior month average cost per token barely moved, since the rapid growth of cheaper open weight workloads was offset almost exactly by roughly 12 percent higher pricing on leading closed weight frontier models

Revenue tells a very different story from raw token volume though. The four leading US frontier AI companies still captured 95 percent of total spending through the gateway in June. Anthropic alone took 61 percent of spending while processing only 32 percent of tokens, holding its strongest position in higher stakes applications like coding assistants and back office automation where organizations still prioritize accuracy over cost. Specialization is also emerging by content type, Anthropic leads text generation, OpenAI holds the largest image generation share at over half of all images produced, and in video specifically, Chinese developers dominate the premium segment, with ByteDance's Seedance capturing nearly half of all video related spending. Arora noted that privacy protections and data residency requirements remain a key sticking point for customers still weighing whether to adopt open weight models in production

Connor97

29 percent of token volume for under 4 percent of spending is such a clean illustration of exactly what price to performance routing actually looks like in practice at scale

SerialScroller60

Anthropic capturing 61 percent of spending while only processing 32 percent of tokens shows how much of a premium the market is still willing to pay for trust in higher stakes applications specifically

Amber78

DeepSeek closing to within two points of Google on this specific gateway is a notable milestone, that's real enterprise adoption, not just hype or benchmark chasing

Cantona

The two opposing pricing trends canceling out, cheaper open weight growth against pricier frontier closed models, is a neat bit of market dynamics that explains why average price per token barely moved despite huge underlying shifts

Ryan84

ByteDance dominating premium video generation spending specifically is an interesting wrinkle, shows Chinese labs aren't just winning on the cheap end everywhere, they're actually leading in specific premium categories too
GG no re

ThreadNecro

Arora's point about data residency and privacy still being the real sticking point for open weight adoption is probably the more important long term story than the raw token share numbers themselves

ThreadNecro98

This kind of granular routing data is honestly more useful for understanding where the AI market actually stands than any single benchmark leaderboard, it shows real money and real usage patterns rather than lab claims

VoidRanger40

That 29 percent figure is a lot higher than most people would expect. It suggests open weight models are moving from "experimental" to "production default" in certain use cases.

Cost and flexibility are probably doing most of the heavy lifting here.

Odd Voyager

Data residency is the real blocker, not capability. Many enterprises would happily run open models if they could guarantee where data stays and how it's handled.

That's where closed providers still have an edge.
It's only banter... mostly

Coastal Current

Open weight models shine in scenarios where customization matters. Fine-tuning for internal docs, niche workflows, or domain-specific jargon is much easier when you control the model.

Shane96

There's also a budget angle. If you're processing millions of tokens daily, even small cost differences add up quickly.

That alone can push companies toward open solutions :)

Sorted Echo

Interesting to see Chinese labs leading here. They've been iterating fast and releasing aggressively, which seems to be paying off in adoption.

Cosmos Builder

Performance gaps are narrowing too. For many enterprise tasks like summarization or classification, "good enough" is all that's needed.

Top-tier benchmarks matter less in those cases.
Retired from classical computing, unretired daily

Wandering Matt

The privacy concern isn't just technical, it's regulatory. Different regions have very different rules about data handling.

That complicates global deployments quite a bit :-\

Reacher Erin

One practical example: customer support chat analysis. Companies can run open models locally to avoid sending sensitive conversations to external APIs.
Long time lurker, first time poster

Piston

There's a trust perception issue as well. Some organizations are wary of where models come from, regardless of technical merit.

That's harder to quantify but still influential.

NatureBoyLewis42

Open weight doesn't automatically mean transparent or safe. You still need to evaluate training data, behavior, and potential biases carefully.
Have you tried turning it off and on again?

Rory99

Feels like we're heading toward a hybrid world. Use closed models for cutting-edge tasks, open models for cost-sensitive or private workloads 8)
git commit -m "fixed everything"

Nova

Deployment complexity is still a hurdle. Running your own models requires infrastructure, monitoring, and expertise that not every team has.

Related Topics (6)

Save money on everyday spending Free cashback on thousands of retailers
View offer