Gemini 3.8 Live
Gemini 3.8 Live is Google's September 15, 2026 default Live API model for low-latency voice agents and real-time dialogue.
Provider
Model family
Google Gemini Live
Full-duplex voice model
Cost tier
Live
Status
Current
Release Sep 15, 2026
Why teams choose it
Migrate Live API clients from gemini-3.1-flash-live-preview to gemini-3.8-live
thinking_level is not supported; omit thinking_config. Proactive audio is permanently on.
Async function calling (behavior: NON_BLOCKING) is the default
Synchronous BLOCKING remains available on this SKU. Affective dialogue is removed.
Google's developer post lists $0.005/min audio in and $0.018/min audio out, with a token-equivalent footnote of $3/$12 per 1M
Backend Search Live and Gemini app surfaces are product rollouts, not this API ID.
Tradeoffs to know
- This is not Gemini 3.8 Flash and not Gemini 3.5 Transcribe. Flash is the text/coding workhorse; Transcribe is speech-to-text only.
- Caching, code execution, file search, structured outputs, image generation, Maps grounding, URL context, and Batch API are documented as unsupported on this card.
When not to use this
- Do not invent a ChatGPT-style mini Live ID. Extended Thinking is a separate model code when you need background reasoning.
- Not ideal for simple tasks where cheaper models in the same lineup are good enough.
- Avoid for latency-sensitive real-time chat when raw response speed outweighs reasoning depth.
Technical specs
- Inputs
- text, image, audio, video
- Outputs
- audio, text
- Capabilities
- full-duplex voice, speech-to-speech, interleaved reasoning, asynchronous function calling, visual grounding, search grounding, multilingual live dialogue, audio generation
- License
- Proprietary API
- Model string
gemini-3-8-live
Benchmarks
{
"api_id": "gemini-3.8-live",
"source": "https://ai.google.dev/gemini-api/docs/models/gemini-3.8-live",
"context_tokens": 131072,
"vendor_reported": true,
"max_output_tokens": 65536,
"audio_input_usd_per_minute": 0.005,
"audio_output_usd_per_minute": 0.018
}Google Gemini Live family lineup
Compare with
Gemini 3.8 Live Extended Thinking
Current · 3.8 Live ET · live-extended-thinking
Gemma 2 27B
Gemma 2 27B is Google’s open-weights Gemma family checkpoint balancing quality and deployability for research and pro…
Gemini 3.1 Pro Preview
Gemini 3.1 Pro Preview is Google's preview Gemini 3-series model for complex multimodal reasoning, software engineeri…
Gemini 3.8 Live FAQ
What is Gemini 3.8 Live?
Gemini 3.8 Live is Google's September 15, 2026 default Live API model for low-latency voice agents and real-time dialogue. The stable model code is gemini-3.8-live. It takes text, images, audio, and video, and returns text and audio, with a 131,072-token input limit and 65,536-token output limit. Google positions it f...
When does Gemini 3.8 Live fit best?
Low-latency voice agents on the Gemini Live API
What should teams watch out for with Gemini 3.8 Live?
This is not Gemini 3.8 Flash and not Gemini 3.5 Transcribe. Flash is the text/coding workhorse; Transcribe is speech-to-text only.
Explore next
Models, tools, and comparisons that connect to this reference.