Cohere Releases 2.4B Visual Small Model: Document Understanding Surpasses Ministral 3 3B

律动BlockBeats
律动BlockBeats|Aug 13, 2026 08:05
According to monitoring by Beating, Cohere has open-sourced its smallest visual language model to date, North Micro Vision, with only 2.4B parameters and licensed under Apache 2.0. It specializes in documents, tables, charts, screenshots, and OCR, capable of directly processing images in their original scale and resolution without needing to resize them to fixed dimensions. The model consists of a 2B language model and a 400M visual encoder. In official tests, its DocVQA score reached 92.1%, surpassing Ministral 3 3B's 89.6% and Gemma 4 E2B's 73.2%, and coming close to Qwen3.5-2B's 92.6%. For visual localization, its RefCOCO score was 73.2%, also significantly higher than Ministral and Gemma. However, it is not the strongest model of its size. Qwen3.5-2B remains superior across most general visual, OCR, and multimodal benchmarks, while North Micro Vision's strengths are more focused on document understanding and visual localization. It is also not a reasoning model, does not support tool invocation, and has limited mathematical and coding capabilities. [Original Link]
+2
Mentioned
Share To

Timeline

HotFlash

APP

X

Telegram

Facebook

Reddit

CopyLink

Hot Reads