律动BlockBeats|9月 02, 2026 03:11
[Qwen3.8-Max Upgraded Again: All 8 Programming Tests Improved, Some Surpassing Opus 5]
Beating AI News Flash: Alibaba's flagship Qwen3.8-Max, released just a month ago, has been upgraded again with the launch of Qwen3.8-Max-0902. The parameter scale remains unchanged, still at 2.4T parameters and 1 million tokens of context. This upgrade primarily focuses on post-training for coding and coworking, emphasizing programming, complex workflows, and long-duration agent tasks.
According to the comparison table released by Alibaba, 0902 outperforms the original version in all 8 programming evaluations. TerminalBench 3.0 improved from 11.3 to 29.0, DeepSWE 1.1 rose from 56.6 to 69.3, and QwenSWEbench V2 increased from 55.1 to 70.0. The professional task benchmark JobBench also climbed from 53.4 to 64.0. Compared to Claude Opus 5, 0902 still falls short in most programming projects but surpasses it in MLS-Bench-Lite, SWE-Atlas QnA, and QwenSWEbench V2.
In multimodal evaluations, the scores only increased by 0.4 to 3 points compared to the original version, clearly indicating that this upgrade focuses on enhancing the agent's coding and task execution capabilities. These results are currently self-tested by Alibaba.
Qwen3.8-Max-0902 is now available on the QwenCloud API, with pricing unchanged at $2 per input and $6 per output per million tokens. [Original Link]
Share To
Timeline
HotFlash
APP
X
Telegram
CopyLink