律动BlockBeats
律动BlockBeats|Aug 11, 2026 15:09
[NVIDIA Open-Sources Nemotron 3.5 Lightning: Designed for Agents to Execute Tool Calls] According to monitoring by Beating, NVIDIA has released the open-source model Nemotron 3.5 Lightning, specifically designed for long-running Agents to handle 'execution tasks.' The model has a total of 30 billion parameters, with only 3 billion activated per token. It primarily focuses on high-frequency tasks such as tool calls, result verification, and sub-Agent scheduling. NVIDIA's approach is to further refine the division of labor among Agents. Complex planning is delegated to larger models like Nemotron 3 Ultra, while repetitive execution is handled by Lightning. According to the official announcement, its output speed can reach up to 4 times that of models in the same class. On PinchBench, Lightning achieved an accuracy rate of 86%, and under similar accuracy conditions, it completed 10,000 tasks 30% faster than Qwen3.6 35B. The model adopts an MoE architecture and has been specifically trained for Agent Harness, incorporating multi-token prediction and speculative decoding. The official release includes BF16 and NVFP4 weights, which can run on local devices such as RTX 5090 and DGX Spark. It also supports llama.cpp, Ollama, LM Studio, and Unsloth. Customization is also emphasized. NVIDIA has made the weights, partial training data, and training recipes available, allowing enterprises to further fine-tune the model for tasks related to code, security, legal matters, and more. In the official demonstrations, Lightning showed significant performance improvements across multiple specialized tasks after fine-tuning. [Original Link]
+3
Mentioned
Share To

Timeline

HotFlash

APP

X

Telegram

Facebook

Reddit

CopyLink

Hot Reads