律动BlockBeats
律动BlockBeats|Sep 09, 2026 04:19
**[Ant Group Announces Open-Source Multimodal Large Model Ling-3.0-flash-VL]** BlockBeats News, September 9: Ant Group announced the official release and open-sourcing of the first native multimodal large model in the Bai Ling series, Ling-3.0-flash-VL. The model is an extension of the Ling-3.0-flash MoE architecture, with a total parameter count of 124B, activating 5.5B parameters per inference. It natively supports image, text, and video inputs, with a context window reaching 256K tokens. Focusing on "how to complete real-world tasks more reliably and efficiently," Ling-3.0-flash-VL explores the following directions: Adding visual capabilities to large models has raised a common concern that it might lower text intelligence. However, training practices have shown the opposite: native multimodal joint training not only expands application boundaries but also enhances text intelligence. Ling-3.0-flash-VL introduces a visual feedback mechanism—observing execution results, comparing them to targets, identifying deviations, and continuously making corrections—transforming tasks from a one-time "generation" process into a closed loop of "observe → act → verify → correct," making execution results more reliable. Ling-3.0-flash-VL inherits the core advantage of Ling-3.0-flash as an efficient execution node in Agent workflows. Within the visual feedback loop, it balances output quality and execution efficiency, advancing complete tasks at lower costs and in shorter timeframes. [Original Link]
+1
Mentioned
Share To

Timeline

HotFlash

APP

X

Telegram

Facebook

Reddit

CopyLink

Hot Reads