GenAIWiki

Gemini 3.5 Transcribe

CurrentLatest

Gemini 3.5 Transcribe is Google's speech-to-text model for precise, formatted transcription of live and recorded audio.

Provider

Google

Model family

Google Gemini Audio

Speech-to-text transcription model

Cost tier

Transcribe

Status

Current

Release Aug 26, 2026

Why teams choose it

🧠

Google publishes two preview endpoints: gemini-3.5-transcribe-live for bidirectional str…

eaming and gemini-3.5-transcribe for recorded audio with timestamps.

📎

Google cites Artificial Analysis WER of 4.0% streaming and 2.6% non-streaming

plus FLEURS multilingual WER of 5.50% / 5.04%.

⚙️

The announcement positions it for voice agents

captioning, and post-call analytics, including filler-word cleanup and custom vocabulary.

Tradeoffs to know

  • Developer and enterprise access is public preview; consumer surfaces and languages vary by product and region.
  • Pre-recorded speaker attribution is documented for up to three speakers; more than three is experimental.
  • WER and latency figures are vendor-reported via Artificial Analysis and should be remeasured on your audio domain.

When not to use this

  • Not ideal for simple tasks where cheaper models in the same lineup are good enough.
  • Avoid for regulated or high-stakes outputs without evaluations that mimic your tooling, data, and review process.
  • Pair catalog notes with comparisons and your own benchmarks before declaring a routing winner.

Technical specs

Inputs
audio
Outputs
text
Capabilities
speech-to-text, streaming transcription, batch transcription, speaker attribution, custom vocabulary, multilingual ASR, disfluency cleanup
License
Proprietary API (public preview)
Model string
gemini-3-5-transcribe

Benchmarks

{
  "source": "https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5-transcribe/",
  "vendor_reported": true,
  "aa_wer_streaming_pct": 4,
  "aa_wer_non_streaming_pct": 2.6,
  "fleurs_wer_streaming_pct": 5.5,
  "fleurs_wer_non_streaming_pct": 5.04
}

Compare with

Gemini 3.5 Transcribe FAQ

What is Gemini 3.5 Transcribe?

Gemini 3.5 Transcribe is Google's speech-to-text model for precise, formatted transcription of live and recorded audio. Google documents public-preview access in the Gemini API, Google AI Studio, Google Antigravity, and Gemini Enterprise Agent Platform, with product surfaces including Rambler on Android and the Gemini...

When does Gemini 3.5 Transcribe fit best?

Live voice agents and captioning

What should teams watch out for with Gemini 3.5 Transcribe?

Developer and enterprise access is public preview; consumer surfaces and languages vary by product and region.

Explore next

Models, tools, and comparisons that connect to this reference.