GenAIWiki

Gemini 3.8 Live

CurrentLatestFrontier

Gemini 3.8 Live is Google's September 15, 2026 default Live API model for low-latency voice agents and real-time dialogue.

Provider

Google

Model family

Google Gemini Live

Full-duplex voice model

Cost tier

Live

Status

Current

Release Sep 15, 2026

Why teams choose it

🧠

Migrate Live API clients from gemini-3.1-flash-live-preview to gemini-3.8-live

thinking_level is not supported; omit thinking_config. Proactive audio is permanently on.

📎

Async function calling (behavior: NON_BLOCKING) is the default

Synchronous BLOCKING remains available on this SKU. Affective dialogue is removed.

⚙️

Google's developer post lists $0.005/min audio in and $0.018/min audio out, with a token-equivalent footnote of $3/$12 per 1M

Backend Search Live and Gemini app surfaces are product rollouts, not this API ID.

Tradeoffs to know

  • This is not Gemini 3.8 Flash and not Gemini 3.5 Transcribe. Flash is the text/coding workhorse; Transcribe is speech-to-text only.
  • Caching, code execution, file search, structured outputs, image generation, Maps grounding, URL context, and Batch API are documented as unsupported on this card.

When not to use this

  • Do not invent a ChatGPT-style mini Live ID. Extended Thinking is a separate model code when you need background reasoning.
  • Not ideal for simple tasks where cheaper models in the same lineup are good enough.
  • Avoid for latency-sensitive real-time chat when raw response speed outweighs reasoning depth.

Technical specs

Inputs
text, image, audio, video
Outputs
audio, text
Capabilities
full-duplex voice, speech-to-speech, interleaved reasoning, asynchronous function calling, visual grounding, search grounding, multilingual live dialogue, audio generation
License
Proprietary API
Model string
gemini-3-8-live

Benchmarks

{
  "api_id": "gemini-3.8-live",
  "source": "https://ai.google.dev/gemini-api/docs/models/gemini-3.8-live",
  "context_tokens": 131072,
  "vendor_reported": true,
  "max_output_tokens": 65536,
  "audio_input_usd_per_minute": 0.005,
  "audio_output_usd_per_minute": 0.018
}

Google Gemini Live family lineup


Compare with

Gemini 3.8 Live FAQ

What is Gemini 3.8 Live?

Gemini 3.8 Live is Google's September 15, 2026 default Live API model for low-latency voice agents and real-time dialogue. The stable model code is gemini-3.8-live. It takes text, images, audio, and video, and returns text and audio, with a 131,072-token input limit and 65,536-token output limit. Google positions it f...

When does Gemini 3.8 Live fit best?

Low-latency voice agents on the Gemini Live API

What should teams watch out for with Gemini 3.8 Live?

This is not Gemini 3.8 Flash and not Gemini 3.5 Transcribe. Flash is the text/coding workhorse; Transcribe is speech-to-text only.

Explore next

Models, tools, and comparisons that connect to this reference.