Gemini 3.5 Transcribe
Gemini 3.5 Transcribe is Google's speech-to-text model for precise, formatted transcription of live and recorded audio.
Provider
Model family
Google Gemini Audio
Speech-to-text transcription model
Cost tier
Transcribe
Status
Current
Release Aug 26, 2026
Why teams choose it
Google publishes two preview endpoints: gemini-3.5-transcribe-live for bidirectional str…
eaming and gemini-3.5-transcribe for recorded audio with timestamps.
Google cites Artificial Analysis WER of 4.0% streaming and 2.6% non-streaming
plus FLEURS multilingual WER of 5.50% / 5.04%.
The announcement positions it for voice agents
captioning, and post-call analytics, including filler-word cleanup and custom vocabulary.
Tradeoffs to know
- Developer and enterprise access is public preview; consumer surfaces and languages vary by product and region.
- Pre-recorded speaker attribution is documented for up to three speakers; more than three is experimental.
- WER and latency figures are vendor-reported via Artificial Analysis and should be remeasured on your audio domain.
When not to use this
- Not ideal for simple tasks where cheaper models in the same lineup are good enough.
- Avoid for regulated or high-stakes outputs without evaluations that mimic your tooling, data, and review process.
- Pair catalog notes with comparisons and your own benchmarks before declaring a routing winner.
Technical specs
- Inputs
- audio
- Outputs
- text
- Capabilities
- speech-to-text, streaming transcription, batch transcription, speaker attribution, custom vocabulary, multilingual ASR, disfluency cleanup
- License
- Proprietary API (public preview)
- Model string
gemini-3-5-transcribe
Benchmarks
{
"source": "https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5-transcribe/",
"vendor_reported": true,
"aa_wer_streaming_pct": 4,
"aa_wer_non_streaming_pct": 2.6,
"fleurs_wer_streaming_pct": 5.5,
"fleurs_wer_non_streaming_pct": 5.04
}Compare with
Gemma 2 27B
Gemma 2 27B is Google’s open-weights Gemma family checkpoint balancing quality and deployability for research and pro…
Gemini 3.1 Pro Preview
Gemini 3.1 Pro Preview is Google's preview Gemini 3-series model for complex multimodal reasoning, software engineeri…
Gemini 2.0 Flash
Gemini 2.0 Flash is Google’s efficiency-oriented multimodal model generation aimed at fast agentic and interactive ex…
Gemini 3.5 Transcribe FAQ
What is Gemini 3.5 Transcribe?
Gemini 3.5 Transcribe is Google's speech-to-text model for precise, formatted transcription of live and recorded audio. Google documents public-preview access in the Gemini API, Google AI Studio, Google Antigravity, and Gemini Enterprise Agent Platform, with product surfaces including Rambler on Android and the Gemini...
When does Gemini 3.5 Transcribe fit best?
Live voice agents and captioning
What should teams watch out for with Gemini 3.5 Transcribe?
Developer and enterprise access is public preview; consumer surfaces and languages vary by product and region.
Explore next
Models, tools, and comparisons that connect to this reference.