律动BlockBeats
律动BlockBeats|Sep 11, 2026 10:33
[Tencent T1 Specializes in Long Agent Tasks: Over 300 Consecutive Operations, Score Increased by 20 Points] Dongcha Beating AI News: Tencent has specifically trained Qwen3.5-122B-A10B for terminal Agent tasks, creating T1. It primarily handles complex tasks in Linux terminals, with a single task capable of consecutively invoking tools over 300 rounds. The Terminal-Bench 2.1 score increased from the base model's 43.8% to 64.0%, a rise of 20.2 points. Of these 20.2 points, the majority came from reinforcement learning. SFT only raised the score to 49.4%, while reinforcement learning added another 14.6 points. The team prepared approximately 15,000 terminal tasks, each equipped with automated testing. The Agent doesn't need to complete the entire task; as long as it fulfills some of the requirements, it can earn corresponding rewards. Long tasks come with another challenge. When the Agent performs tasks and later uses the experience for training, the same content might be re-segmented into different tokens, and MoE might switch to a different set of experts for processing. T1 records the tokens generated at the time and the selected experts, replaying them directly during training, reducing this training-inference discrepancy by about one-third. [Original Link]
+5
Mentioned
Share To

Timeline

HotFlash

APP

X

Telegram

Facebook

Reddit

CopyLink

Hot Reads