Muse Voice Transcribe
Muse Voice Transcribe is Meta Superintelligence Labs' September 1, 2026 real-time audio perception model.
Provider
Meta
Model family
Meta Muse Audio
Speech-to-text transcription model
Cost tier
Transcribe
Status
Current
Release Sep 1, 2026
Why teams choose it
Meta bills $0.18 per hour of audio processed for streaming and file transcription
There is no contributor/training-discount tier at launch.
Rate limits are 8 concurrent streams and 1
000 streams per hour, not token RPM/TPM.
The model returns transcript text only: no speech synthesis
no word-level timestamps, and no emotion or sound-event labels.
Tradeoffs to know
- Meta recommends 25 extensively verified languages at launch even though training covers 70+.
- It is not a speech-to-speech or TTS API.
- Vendor Artificial Analysis ranking claims should be remeasured on your audio domain.
When not to use this
- Not ideal for simple tasks where cheaper models in the same lineup are good enough.
- Avoid for regulated or high-stakes outputs without evaluations that mimic your tooling, data, and review process.
- Pair catalog notes with comparisons and your own benchmarks before declaring a routing winner.
Technical specs
- Inputs
- audio
- Outputs
- text
- Capabilities
- speech-to-text, streaming transcription, file transcription, speaker diarization, endpointing, code-switching, keyword biasing, multilingual ASR
- License
- Proprietary API
- Model string
muse-voice-transcribe
Benchmarks
{
"api_id": "muse-voice-transcribe-1.0",
"source": "https://research.meta.ai/blog/introducing-muse-voice-transcribe",
"vendor_reported": true
}Compare with
Meta
Llama 3 70B
Catalog entry for this named release; see the provider’s official documentation for modalities, pricing, and context…
Meta
Llama 3 8B
Catalog entry for this named release; see the provider’s official documentation for modalities, pricing, and context…
Meta
Llama 3.1 8B Instruct
Llama 3.1 8B Instruct is a small open-weights model for edge laptops, single-GPU servers, and ultra-low-latency assis…
Muse Voice Transcribe FAQ
What is Muse Voice Transcribe?
Muse Voice Transcribe is Meta Superintelligence Labs' September 1, 2026 real-time audio perception model. The API ID is muse-voice-transcribe-1.0. It does streaming ASR with adaptive delay, speaker diarization for 20+ speakers, and native endpointing. Access is via Meta Model API (WebSocket realtime and file POST /v1/...
When does Muse Voice Transcribe fit best?
Live voice agents and captioning
What should teams watch out for with Muse Voice Transcribe?
Meta recommends 25 extensively verified languages at launch even though training covers 70+.
Explore next
Models, tools, and comparisons that connect to this reference.