律动BlockBeats|Jul 25, 2026 03:19
[AMD Releases Fully Open-Source MoE Large Model Instella-MoE, 16B Parameters to Challenge Mainstream Open-Source Models]
BlockBeats News, July 25: AMD announced the launch of its fully open-source Mixture of Experts (MoE) model, Instella-MoE, which features a total of 16 billion parameters, with 2.8 billion parameters activated per token. It is claimed to deliver leading performance among open-source language models of the same scale.
AMD stated that Instella-MoE was entirely trained from scratch using its own AMD Instinct MI300X and MI325X GPUs, along with the ROCm software stack. The model incorporates architectural innovations such as Gated Multi-head Latent Attention (Gated MLA) and FarSkip-Collective to enhance training and inference efficiency.
Performance tests show that the Instella-MoE-16B-A3B base model achieved an average score of 76.7, ranking among the top open-source models, surpassing models like SmolLM3-3B and OLMo-3-7B. Additionally, it competes with larger-scale models while activating only 2.8 billion parameters. The model also supports long-context processing of up to 64K tokens and has completed the full training pipeline, including pre-training, intermediate training, long-context extension, supervised fine-tuning (SFT), direct preference optimization (DPO), and reinforcement learning (RL).
AMD has simultaneously released all model weights, training configurations, data ratios, intermediate checkpoints, and inference code for Instella-MoE, aiming to promote open AI model research and reproducibility. AMD emphasized that Instella-MoE demonstrates the capability of training large-scale MoE models based on AMD hardware and open software ecosystems. The company plans to continue advancing the development of larger-scale, more powerful, and more efficient open-source language models in the future. [Original Article Link]
Share To
Timeline
HotFlash
APP
X
Telegram
CopyLink