DeepSeek Makes 75 Percent V4-Pro Price Cut Permanent: Output Tokens Now $0.87 Per Million

Started by NeutrinoX74, Jun 27, 2026, 06:26 AM

Previous topic - Next topic

0 Members and 1 Guest are viewing this topic.

Topic: DeepSeek Makes 75 Percent V4-Pro Price Cut Permanent: Output Tokens Now $0.87 Per Million   Views(Read 47 times)

NeutrinoX74

DeepSeek announced in May that the 75 percent promotional price cut on its V4-Pro API, originally set to expire May 31, would become the permanent list price. Output tokens now cost 0.87 dollars per million and input tokens range from 0.003625 to 0.435 dollars per million depending on caching. For context, GPT-5.5 charges 2.50 dollars per million input and 10 dollars per million output. Anthropic's Claude Opus 4.7 sits at 5 dollars input and 25 dollars output. DeepSeek's permanent rate is 11 times cheaper on output than GPT-5.5 and 29 times cheaper than Claude Opus 4.7.

The reason the company could make this cut permanent rather than reverting to launch pricing is architectural. V4-Pro was engineered specifically to run efficiently on Huawei Ascend hardware using a Mixture-of-Experts design that activates only a fraction of its parameters during inference. At 1 million token context length it reportedly runs at roughly 27 percent of the per-token compute and 10 percent of the memory of its predecessor. The price cut is an efficiency gain being passed through to the API, not a margin sacrifice.

The implications for the AI pricing market are significant and already playing out. The era of high-margin frontier model tokens that US labs enjoyed through 2025 is under structural pressure. Google has repeatedly cut Gemini prices to compete. OpenAI's pivot toward consumer platform features including advertising reflects the reality that pure API token revenue may not sustain trillion dollar valuations. Enterprise buyers now have a credible, capable alternative for high-volume inference workloads at a fraction of Western model costs, with the significant caveat of data sovereignty and geopolitical risk.


NeonPhantom39

The cache-hit pricing at 0.003625 per million is the number that matters most for agentic workloads. Agent loops re-send system prompts constantly. At that price long-running agents become economically viable at scales that were impossible six months ago

Undertaker

The data sovereignty concern is real and not dismissible. Routing sensitive enterprise data through a Chinese API provider in 2026, given everything happening between the US and China, is a meaningful risk that the price discount has to be weighed against
Be excellent to each other

RomanReigns02

GPT-5.5 is genuinely 11 times more expensive on output. For commodity inference workloads where top percentile reasoning quality is not required that gap is very hard to justify commercially. Enterprise procurement teams are going to have difficult conversations

Dank

The Huawei Ascend dependency is interesting. If the Ascend 950 supply ramp goes well DeepSeek gets even cheaper to run. If it hits supply constraints they have a problem. This is not running on universally available hardware

Dom9

I keep hearing that DeepSeek is within 3 to 7 percentage points of GPT-5.5 on most coding benchmarks. If that is accurate and the price difference is 11x then for a huge number of developer use cases the migration math is obvious

NightOwl94

Anthropic being partially banned and DeepSeek being permanently cheaper is a difficult combination for Western AI labs. The market share data showing Claude at 10.3 percent may look very different in three months
Not financial advice. Not medical advice. Just vibes.

BigDog_Fan

The V4-Flash variant is even cheaper. 0.14 input and 0.28 output per million. That is below the cost of running most open source models on your own compute at reasonable scale. The economics of this keep getting harder to argue against

Related Topics (4)

Save money on everyday spending Free cashback on thousands of retailers
View offer