DeepSeek V4 Pro Midnight Raid: Agent's Score Equaled Claude Fable 5 Overnight, the Entry Fee for the Flagship Intelligence Agent Dropped to Chump Change

CN
1 hour ago
The entry fee for flagship intelligent agents has been breached again for developers and AI entrepreneurs.

Author: Claude, Deep Tide TechFlow

Deep Tide Overview: Early on August 13, DeepSeek unexpectedly launched the official version of V4 Pro: the same API name, yet the Agent capability jumped from the preview version's "hardly usable" to nearly competing with Anthropic's flagship Claude Fable 5, with some rankings achieving a turnaround. Even more aggressive is the price, which is only a fraction of overseas comparable models. For developers and AI entrepreneurs, the entry fee for flagship intelligent agents has been breached again. The only issue is that the official price increase notice has already been posted.

In the early hours of today, the DeepSeek official website was quietly updated, with the DeepSeek-V4-Pro model version upgraded to the formal version 0813; developers do not need to change any code, simply calling deepseek-v4-pro will access the latest model. The new model supports a 1 million token context and up to 384,000 tokens output, with the thinking mode enabled by default, and is compatible with OpenAI, Anthropic, and Responses API formats. Mainstream Agent tools like Codex, Claude Code, and OpenCode can be directly integrated. The news also reached the front page of Hacker News (over 700 upvotes) and the trending list on Zhihu (over 20 million热度), becoming the biggest consensus hotspot in today's Chinese and English AI circles.

Surprise announcement: The same API name can now perform tasks it couldn't yesterday

The core of this update is the Agent capability. According to the official evaluation, the V4 Pro formal version's score on its own software engineering benchmark DeepSWE skyrocketed from 12.8 in the preview version to 62.7; the natural language generation code repository NL2Repo increased from 38.5 to 61.5; the full stack development difficult subset DSBench-Hard doubled to 67.2.

A domestic technology media’s practical test stated: "The same API name, tasks that it basically couldn't accomplish yesterday can now compete with the top tier code Agents globally." In a practical test by aifaner, V4 Pro could autonomously decompose vague instructions like "collect recent tech news to create a mock traditional media homepage," executing it step-by-step and achieving a high degree of completion. Developers on Hacker News also shared their real-world bill: using it for a full day on traffic simulation and distributed physics engine optimization, 2 billion tokens cost about $12.5, "achieving significant performance improvements without introducing any new issues."

Security offense and defense surpass Fable 5, but it is not yet at the top of the open-source hierarchy

Officially, V4 Pro scored 83.3 on the AI security offense and defense benchmark Cybergym, surpassing Claude Fable 5's 83.1; the workflow intelligent agent AutomationBench also achieved surpassing results; high difficulty reasoning HLE (with tools enabled) scored 60.0, just behind Fable 5's 63.0; terminal operation Terminal Bench 2.1 scored 87.9, higher than Claude Opus 4.8's 85.0.

However, two points need to temper enthusiasm. First, it still has a slight gap with Kimi K3 of the Dark Side of the Moon (Terminal Bench 2.1's 87.9 compared to 88.3), so the top spot in the open-source world hasn't changed hands yet. Second, third-party evaluations are not as optimistic as the official assessments: OpenRouter cited the Comprehensive Intelligence Index from Artificial Analysis scoring it 45.3, "better than 70% of models," revealing a temperature difference with the official narrative of "equaling Fable 5." The official scores were measured under its Harness framework, and independent verification will take time.

3 yuan versus 10 dollars, the "penny pricing" of flagship Agents

What truly unsettles competitors is the price. The V4 Pro official version maintains the preview version pricing: 0.025 yuan per million tokens for cache hit input, 3 yuan for cache miss input, and 6 yuan for output. In comparison, Kimi K3's API price is $3 for input and $15 for output; Claude Fable 5 is $10 for input and $50 for output. Even considering currency differences, DeepSeek is only a fraction of overseas comparable models.

Data from OpenRouter shows that due to a cache hit rate of 86%, users' actual weighted input price is as low as $0.064 per million tokens. To borrow a saying from the domestic community: DeepSeek's order remains "those stronger than me aren’t cheaper, those cheaper than me aren’t stronger."

Price increase notice has been posted, API output "stiff language" criticized

Low prices may be time-limited. DeepSeek announced on August 6 that it would "soon adjust API pricing overall, with a significant increase," and plans to introduce peak and valley pricing: during peak times from 9 AM to 12 PM and 2 PM to 6 PM Beijing time, V4 Pro cache miss input will rise from 3 yuan to 6 yuan, and output from 6 yuan to 12 yuan, directly doubling. Since the official version release, prices have not yet been adjusted, allowing developers a well-understood window of time. Another countering piece of information is that DeepSeek promised during the preview version release that after the mass market launch of the Ascend 950 super nodes in the second half of the year, the Pro price will be significantly lowered. The gap between increases and decreases depends on the arrival pace of domestic computing power.

Additionally, developers on X have reported significant differences in performance between the V4 Pro web version and API calls, stating the API output language "reads stiffly and with awkward syntax," to which DeepSeek has not yet responded. Teams relying on the API for products should test it themselves before going live.

After the cost collapse of flagship Agents, the next focus is Harness

Put this matter into a larger picture:

This week, a discussion was sparked by research revealing "inference traces stolen from closed-source model APIs," indicating cracks in the reasoning moat of closed-source vendors. Now DeepSeek proves, with "penny pricing plus first-tier scores," that the capability gap itself is also being narrowed. The remaining pricing power of closed-source flagships has loosened a bit.

For AI application entrepreneurs and investors focusing on the computing power chain, deeper clues are hidden in the official documentation. DeepSeek revealed in the V4 Flash update log that its Agent operating framework DeepSeek Harness is about to be released, with recruitment and official public accounts already in place. The industry consensus is that when model inference capabilities approach a ceiling, the system that decides the Agent's upper limits will be something like Harness. With the model in place, if the framework delivers, the next DeepSeek moment won’t be far off.

免责声明:本文章仅代表作者个人观点,不代表本平台的立场和观点。本文章仅供信息分享,不构成对任何人的任何投资建议。用户与作者之间的任何争议,与本平台无关。如网页中刊载的文章或图片涉及侵权,请提供相关的权利证明和身份证明发送邮件到support@aicoin.com,本平台相关工作人员将会进行核查。

Share To
APP

X

Telegram

Facebook

Reddit

CopyLink