PANews|9月 11, 2026 06:22
[Xiaomi Releases and Open-Sources Xiaomi-CocktailASR-1]
Xiaomi has officially released and open-sourced the industrial-grade target speaker speech recognition model, Xiaomi-CocktailASR-1. According to the introduction, Xiaomi-CocktailASR-1 is designed to address the 'cocktail party' problem. The model adopts an end-to-end LLM architecture, using a reference audio clip of the target speaker as a voiceprint prompt. It can accurately extract and transcribe only the target user's speech in complex environments where multiple people are speaking simultaneously.
Share To
Timeline
HotFlash
APP
X
Telegram
CopyLink