DeepSeek Prices Its New V4-Pro-0813 Model At $0.87 Per 1 Million Output Tokens, As The High-Flying Chinese AI Lab Wows With Its Soaring Token Consumption

DeepSeek was second only to Anthropic in terms of the total number of tokens consumed in July. And now, perhaps in a bid to cement its ascendancy, the high-flying Chinese AI lab has just unveiled the DeepSeek-V4-Pro-0813 model, its latest gambit to take on the might of OpenAI and Anthropic.

DeepSeek has started rolling out the V4-Pro on its API and Chat. The AI lab has priced the model at $0.435 per 1 million tokens of input and $0.87 per 1 million tokens of output.

Do note that OpenAI launched a literal price war a few days back by discounting its GPT-5.6 Luna by as much as 80 percent , with input tokens now priced at just $0.2 per 1 million from their earlier perch at $1, and output tokens priced at just $1.20 per 1 million vs. the earlier price of $6.

Just hours later, however, DeepSeek launched a refreshed version of its latest Flash-class model, dubbed the V4-Flash-0731. Critically, the model has just 284 billion parameters and yet offers a performance that is similar to Anthropic's Opus 4.8, which is widely believed to span multi-trillion parameters!

And, in what went right to the heart of OpenAI's price war, DeepSeek priced the V4 Flash 0731 at just $0.14 per 1 million tokens of input , and $0.28 per 1 million tokens of output, eviscerating any comparative price advantage that OpenAI tried to garner with its discounting move.

Coming back, DeepSeek's V4-Pro model outcompetes Opus 4.8 on Terminal Bench 2.1, Cybergym, DeepSWE, and AutomationBench benchmarks, as per the preliminary results populating WeChat right now.

Meanwhile, as stated earlier, DeepSeek was second only to Anthropic in terms of token volume in July, and might even clinch the apex spot in the coming months.

As such, DeepSeek is currently contending with an unprecedented demand surge, especially amid anecdotes that suggest its models' inference speeds slow down to a crawl at times, which is wholly understandable given the lab's limited compute footprint of just around 20,000 NVIDIA H100 GPUs .

Follow Wccftech on Google to get more of our news coverage in your feeds.