Late on the night of August 12, DeepSeek’s official API documentation was updated, with the model version switching from the V4-Pro preview to “DeepSeek-V4-Pro-0813.” Almost at the same time, Elon Musk’s SpaceXAI released its next-generation model, Grok 4.6.
The two cost-effective models collided on the same night. In the face-off between Liang Wenfeng and Musk, the companies that may be feeling the real pressure are OpenAI and Anthropic.
An evaluation comparison table circulating from DeepSeek’s official user community shows that the V4 Pro final version performed impressively across multiple agent benchmarks. On Terminal Bench 2.1, which measures an AI agent’s ability to complete complex tasks in real-world terminal environments, V4 Pro scored 87.9, just 0.1 point behind Anthropic’s Fable 5, the global leader. Compared with the preview version’s score of 72.1, that represents a 15.8-point jump in just three months.
On Cybergym, a benchmark for AI security agents, V4 Pro scored 83.3, edging out Fable 5 at 83.1. It also came out ahead on AutomationBench, a workflow-agent benchmark, with 31.8 points versus 29.1.
On DeepSWE, which measures software-engineering capabilities, the score soared from 12.8 in the preview version to 62.7—nearly five times higher—and surpassed Anthropic’s previous flagship, Opus 4.8, which scored 58.0. The final version of V4 Pro supports a context window of 1 million tokens and a maximum output of 384,000 tokens. It is also compatible with both OpenAI and Anthropic API formats, allowing developers to switch over with virtually no code changes.
Agents are the key to understanding this upgrade. Tests such as Terminal Bench and DeepSWE examine whether AI can actually get things done—calling tools, executing multi-step tasks, writing and modifying code, and steadily pushing through long chains of work until a deliverable is produced. This is precisely the threshold that separates large language models as conversational toys from genuine production tools.
Just hours before V4 Pro went live, Grok 4.6, released by Musk’s SpaceXAI, was likewise aimed at long-horizon agentic tasks. In GDPVal-AA v2, an evaluation designed around real-world knowledge work, Grok 4.6 topped the leaderboard with an Elo rating of 1,753, ahead of Fable 5 at 1,741 and GPT-5.6 Sol Max at 1,728. The two models, independently but almost in lockstep, are targeting the same goal: enabling AI to reliably deliver usable results on long-running tasks.
Yet the price gap is staggering. Fable 5 costs $50 per million output tokens, GPT-5.6 Sol Max costs $30, Grok 4.6 costs $6, while DeepSeek V4 Pro costs just $0.87—roughly one fifty-seventh the price of Fable 5. Behind a mere 0.1-point difference in benchmark scores lies a 57-fold gulf in price.
Another major highlight of DeepSeek’s release is the progress of its agentic tool, “Harness.” Media reports say DeepSeek Harness was initiated internally in May 2026, with Cui Tianyi taking charge of the project.
Cui Tianyi graduated from Zhejiang University’s Department of Computer Science and previously spent nine years at quantitative trading firm Jane Street. He joined DeepSeek in March 2026. The “DeepSeek Harness Team” official account was formally registered on July 6 this year. On August 1, Cui publicly solicited beta testers for Harness from around the world.
Alongside the surge in performance came a warning of higher prices. On August 6, DeepSeek announced on its open platform that it planned to raise API service prices across the board in the near future, with a relatively substantial increase expected. This marks the first time DeepSeek has announced an across-the-board API price increase, rather than the partial adjustments it had previously made for peak usage periods.
A look at its pricing timeline is revealing. When V4 was released in April, the V4 Pro API launched at a 75% discount. After the promotion ended on May 31, prices were permanently cut to one-quarter of the original rate. On June 29, DeepSeek introduced peak/off-peak pricing. It took just over two months to go from a permanent price cut to a warning of a major price increase.
If the price hike is understood simply as DeepSeek being “unable to absorb the costs,” its intentions may be underestimated. Just two months ago, on June 16, DeepSeek completed its first-ever external financing round since its establishment, raising more than RMB 50 billion. Its post-money valuation surpassed RMB 350 billion, setting a record for a single financing round in China’s AI industry.
The latest reports indicate that DeepSeek is already pushing ahead with a second round of financing, aiming to raise around RMB 50 billion at a pre-money valuation of approximately RMB 500 billion. A Morgan Stanley research report published on August 9 identified three drivers: strong demand for the V4 model is supporting greater pricing power; AI companies need to strike a balance between market share and gross margins in order to sustain investment in frontier-model R&D; and greater use of domestically produced chips could result in higher inference costs.
The report argues that model intelligence, rather than price, will be the ultimate moat in the competition among large language models.
DeepSeek’s price increase is not an isolated case. According to Morgan Stanley, using ByteDance, Alibaba, Baidu, Tencent, Zhipu, Moonshot AI, MiniMax, and DeepSeek as a sample, the average API output price for Chinese large models has risen from approximately RMB 12.2 per million tokens in the first quarter of 2025 to RMB 21.9 per million tokens in the second quarter of 2026. Zhipu’s API prices have risen by approximately 83% cumulatively since the end of last year, yet its usage volume has increased by 400%. Tencent raised model API prices twice between March and April, with increases reaching as high as 463% for some services.
Across the AI industry, the supply-demand dynamics of computing power, chip procurement costs, and operating expenses are all continuing to rise. A business model based simply on using low prices to buy market share is becoming increasingly difficult to sustain. As leading players move away from treating price as their sole competitive weapon, the focus of the industry’s second half is becoming clear: it is no longer about who is cheaper, but who can actually solve problems.
Editor: Zhiyu Wang



