PANews|Sep 02, 2026 01:08
[Meta Releases Real-Time Audio Perception AI Model Muse Voice Transcribe, Supporting Multi-Speaker Transcription for Over 20 People]
According to IT Home citing 9to5Mac, Meta has launched its first real-time audio perception model, Muse Voice Transcribe. Developers can access it via the Meta Model API, priced at $3 per 1,000 minutes of audio (approximately $0.18 per hour). This model integrates streaming automatic speech recognition, speaker separation, and endpoint detection within the same process. It can output text synchronously as users speak, eliminating the need to wait for the complete recording to finish before processing. The model supports over 70 languages (25 of which have been verified), audio exceeding one hour in length, and multilingual switching. It can separate different speakers in recordings involving more than 20 participants. Meta has introduced an adaptive latency mechanism, enabling rapid output for easily recognizable speech while leveraging more context for complex terms to balance speed and accuracy.
Share To
Timeline
HotFlash
APP
X
Telegram
CopyLink