Google
Latest models
Newest verified model in each family. Spec lines appear only when the model card already stores them.
Gemini 3.1 Flash-Lite is Google's stable Gemini 3-series workhorse model for cost-efficient, high-volume multimodal workloads.
Gemini 3.8 Live Extended Thinking is Google's September 15, 2026 high-reasoning Live API model for complex multi-step work during a live voice session. The stable model code is gemini-3.8-live-extended-thinking. Same 131,072 / 65,536 token limits and text/image/audio/video in, text and audio out as Gemini 3.8 Live. It runs background reasoning and asynchronous tool calls while streaming audio. Google cites #1 on Artificial Analysis Speech-to-Speech Quality Index (82.6), 68.6% on τ-Voice, 35.1% on Sierra τ-Voice-banking, and 97.7% on Big Bench Audio. Same listed audio session rates as 3.8 Live.
API ID gemini-3.8-live-extended-thinking
- Context tokens
- 131,072
- Max output tokens
- 65,536
- Audio input (USD / min)
- 0.005
- Audio output (USD / min)
- 0.018
Gemini 3.8 Flash Cyber is Google's September 2, 2026 defensive-security variant of Gemini 3.8 Flash, gated through the Fairwind Program for trusted government authorities, critical-infrastructure operators, and software maintainers. Google positions it for autonomous vulnerability discovery and automated patching, not public Gemini API self-serve.
Gemini 3.8 Live is Google's September 15, 2026 default Live API model for low-latency voice agents and real-time dialogue. The stable model code is gemini-3.8-live. It takes text, images, audio, and video, and returns text and audio, with a 131,072-token input limit and 65,536-token output limit. Google positions it for scale and cost efficiency: interleaved reasoning, asynchronous function calling, 97-language mid-conversation switching, and visual grounding while the session keeps talking. Audio sessions are listed at $0.005 per minute input and $0.018 per minute output on the developer Live post.
API ID gemini-3.8-live
- Context tokens
- 131,072
- Max output tokens
- 65,536
- Audio input (USD / min)
- 0.005
- Audio output (USD / min)
- 0.018
Gemini 3.8 Flash is Google's September 2, 2026 workhorse for software engineering, agentic tasks, and multi-step reasoning at Flash speed. The DeepMind model card documents text, image, audio, and video input, a 1M-token context window, 64K output, and adjustable effort levels. Google keeps the same introductory price as 3.7 Flash: $0.75 input and $3.75 output per million tokens through December 31, 2026.
- Context tokens
- 1,000,000
- Max output tokens
- 64,000
Gemini 3.5 Transcribe is Google's speech-to-text model for precise, formatted transcription of live and recorded audio. Google documents public-preview access in the Gemini API, Google AI Studio, Google Antigravity, and Gemini Enterprise Agent Platform, with product surfaces including Rambler on Android and the Gemini app on macOS.
WeatherNext 3 is Google DeepMind and Google Research's September 3, 2026 global weather AI model. It ingests live geostationary satellite mosaics and station observations to produce hourly forecasts, with surface variables down to 5 km and atmospheric variables at 25 km. Google is integrating it into Search, the Gemini app, Maps, the Maps Platform Weather API, Earth Engine, BigQuery, and Cloud Storage.
TimesFM-3 is Google Research's 330-million-parameter time-series foundation model for zero-shot univariate and multivariate forecasting. The August 31, 2026 announcement documents native multivariate forecasting in a single forward pass, past and past-future covariates, a pretraining corpus of more than 1 trillion time points, and availability on GitHub and Hugging Face, with BigQuery integration planned in the coming weeks.
Gemma 2 27B is Google’s open-weights Gemma family checkpoint balancing quality and deployability for research and product teams that need permissive terms without Vertex-only APIs. It is often fine-tuned for domain tasks on TPU or GPU clusters.
Gemini Omni 1.1 Flash is Google's production video generation and editing model for developers. The August 27, 2026 announcement documents scene extension from up to 10 seconds of prior context to a 40-second cumulative clip, first-and-last-frame interpolation, 360p drafts, upscaling to 1080p or 4K, and video references, with the API model ID gemini-omni-1.1-flash.
Also current
Other verified current models from Google. Older rows stay in the models directory.