Tech & AI News
Engadget

Meta's new AI transcription model can distinguish between multiple speakers and languages in real-time

Meta introduced Muse Voice Transcribe, its first real-time audio model, which can diarize over 20 speakers, handle simultaneous multilingual speech, and support mid-sentence code-switching across 70+ languages (25 validated at launch). The system uses adaptive delay to predict tokens, maintains accuracy on noisy audio, and is available in the Meta AI Mac app, via Muse Code, and through the Model API at $3 per 1,000 minutes.