律动BlockBeats|Aug 05, 2026 10:43
[Qwen Matches Opus 5 in 3-Round Recall Rate, at Half the Cost]
According to monitoring by Beating, the security company Aikido used its code auditing Agent to test seven models, including Qwen3.8-Max, Claude Opus 5, Kimi K3 Max, DeepSeek V4 Flash, as well as GPT-5.6 Sol, Luna, and Terra. The test involved 32 recently disclosed vulnerabilities, with each model running three times. Qwen3.8-Max identified a total of 26 vulnerabilities across three rounds, achieving a recall rate of 81.3%, tying for first place with Opus 5. Its F1 composite score (balancing false negatives and false positives) was 83.2%, slightly lower than Opus 5, Kimi K3 Max, and GPT-5.6 Sol.
The issue with Qwen lies in its inconsistent performance across individual runs. Out of the 26 vulnerabilities, only 10 were consistently identified in all three rounds, whereas Opus 5 and Sol each identified 19. However, Qwen's total cost was approximately $821, only half that of Opus 5 and Sol, though still five times that of DeepSeek V4 Flash. [Original Link]
Share To
Timeline
HotFlash
APP
X
Telegram
CopyLink