DeepSeek’s V4 Flash Guts OpenAI’s Price War Within Hours, Matching Opus-Grade Output At $0.28 Per Million Tokens, As Moonshot Scales With 20,000 New NVIDIA GPUs

We told our readers not to discount China's uncanny ability to trounce the West in a price war, and that's exactly what its AI labs have done: merely hours after OpenAI fired its first volley, DeepSeek's revamped V4 Flash model now offers unprecedented economy and performance, while the progenitor of the Kimi K3 model, Moonshot, has just secured a sizable cluster of NVIDIA GPUs to crank up its training workloads.

As we detailed recently, OpenAI has just launched a literal price war by discounting its GPT-5.6 Luna by as much as 80 percent , with input tokens now priced at just $0.2 per 1 million from their earlier perch at $1, and output tokens priced at just $1.20 per 1 million vs. the earlier price of $6.

Of course, OpenAI claimed at the time that it was able to implement this steep discount after extracting additional architectural efficiencies from its models. Even so, most interpreted the move as its opening gambit in a price war aimed at China's AI labs.

Just hours later, however, DeepSeek has launched a refreshed version of its latest Flash-class model, dubbed the V4 Flash 0731. Critically, the model has just 284 billion parameters and yet offers a performance that is similar to Anthropic's Opus 4.8, which is widely believed to span multi-trillion parameters!

That's not all. In what goes right to the heart of OpenAI's price war, DeepSeek has priced the V4 Flash 0731 at just $0.14 per 1 million tokens of input , and $0.28 per 1 million tokens of output, eviscerating any comparative price advantage that OpenAI tried to garner with its discounting move.

Meanwhile, Bloomberg has reported separately that Moonshot has secured a compute cluster consisting of 20,000 H200 NVIDIA GPUs from Alibaba, significantly upping its training capacity.

Of course, Moonshot jumped right into the ongoing multi-dimensional struggle between China and the US recently, when Anthropic and some members of the Trump administration accused Moonshot of distilling its Kimi K3 model from Anthropic's Fable.

The US has long suspected that Chinese engineers often take their model-laden hard drives to US-friendly countries to imbue their models with frontier-level capabilities by using reinforcement learning techniques on cutting-edge NVIDIA GPUs.

Apparently, as per the allegations leveled by a Trump administration official recently, Moonshot not only furtively owns some NVIDIA GB300 servers, but was also able to access additional GB300s via Thailand.

Thinking Machine has just unveiled Inkling-Small, an open-weight model that sports just 276 billion parameters. On the Artificial Analysis Intelligence Index, DeepSeek V4 Flash 0731 leads with a score of 50 percent , while Inkling-Small trails slightly at 40 percent, matching the older baseline DeepSeek V4 Flash variant.

Follow Wccftech on Google to get more of our news coverage in your feeds.