Structured cards
Model database
Filter by provider, architecture family, or full-text search across descriptions.
Frontier models
Verified flagships and this month’s launches—Claude Opus 5.5, GPT-6 Sol and Luna, GPT-6 Astra, Claude Fable 5.1, Gemini 3.8, and more. Scroll sideways for the full shelf.
By company
Latest verified models from each frontier company, summarized on one page.
All models
Filter and paginate the full catalog. Tabs control lifecycle scope.
Gemini 3.8 Live
CurrentLatestGemini 3.8 Live is Google's September 15, 2026 default Live API model for low-latency voice agents and real-time dialogue. The stable model code is gemini-3.8-live. It takes text, images, audio, and video, and returns text and audio, with a 131,072-token input limit and 65,536-token output limit. Google positions it for scale and cost efficiency: interleaved reasoning, asynchronous function calling, 97-language mid-conversation switching, and visual grounding while the session keeps talking. Audio sessions are listed at $0.005 per minute input and $0.018 per minute output on the developer Live post.
Gemini 3.8 Live Extended Thinking
CurrentLatestGemini 3.8 Live Extended Thinking is Google's September 15, 2026 high-reasoning Live API model for complex multi-step work during a live voice session. The stable model code is gemini-3.8-live-extended-thinking. Same 131,072 / 65,536 token limits and text/image/audio/video in, text and audio out as Gemini 3.8 Live. It runs background reasoning and asynchronous tool calls while streaming audio. Google cites #1 on Artificial Analysis Speech-to-Speech Quality Index (82.6), 68.6% on τ-Voice, 35.1% on Sierra τ-Voice-banking, and 97.7% on Big Bench Audio. Same listed audio session rates as 3.8 Live.
Gemini 3.8 Flash
CurrentLatestGemini 3.8 Flash is Google's September 2, 2026 workhorse for software engineering, agentic tasks, and multi-step reasoning at Flash speed. The DeepMind model card documents text, image, audio, and video input, a 1M-token context window, 64K output, and adjustable effort levels. Google keeps the same introductory price as 3.7 Flash: $0.75 input and $3.75 output per million tokens through December 31, 2026.
Gemini 3.5 Transcribe
CurrentLatestGemini 3.5 Transcribe is Google's speech-to-text model for precise, formatted transcription of live and recorded audio. Google documents public-preview access in the Gemini API, Google AI Studio, Google Antigravity, and Gemini Enterprise Agent Platform, with product surfaces including Rambler on Android and the Gemini app on macOS.
Gemini Omni 1.1 Flash
CurrentLatestGemini Omni 1.1 Flash is Google's production video generation and editing model for developers. The August 27, 2026 announcement documents scene extension from up to 10 seconds of prior context to a 40-second cumulative clip, first-and-last-frame interpolation, 360p drafts, upscaling to 1080p or 4K, and video references, with the API model ID gemini-omni-1.1-flash.
WeatherNext 3
CurrentLatestWeatherNext 3 is Google DeepMind and Google Research's September 3, 2026 global weather AI model. It ingests live geostationary satellite mosaics and station observations to produce hourly forecasts, with surface variables down to 5 km and atmospheric variables at 25 km. Google is integrating it into Search, the Gemini app, Maps, the Maps Platform Weather API, Earth Engine, BigQuery, and Cloud Storage.
Gemini 3.8 Flash Cyber
CurrentLatestGemini 3.8 Flash Cyber is Google's September 2, 2026 defensive-security variant of Gemini 3.8 Flash, gated through the Fairwind Program for trusted government authorities, critical-infrastructure operators, and software maintainers. Google positions it for autonomous vulnerability discovery and automated patching, not public Gemini API self-serve.
TimesFM-3
CurrentLatestTimesFM-3 is Google Research's 330-million-parameter time-series foundation model for zero-shot univariate and multivariate forecasting. The August 31, 2026 announcement documents native multivariate forecasting in a single forward pass, past and past-future covariates, a pretraining corpus of more than 1 trillion time points, and availability on GitHub and Hugging Face, with BigQuery integration planned in the coming weeks.
Gemini 3.1 Flash-Lite
CurrentLatestGemini 3.1 Flash-Lite is Google's stable Gemini 3-series workhorse model for cost-efficient, high-volume multimodal workloads.
Gemma 2 27B
CurrentLatestGemma 2 27B is Google’s open-weights Gemma family checkpoint balancing quality and deployability for research and product teams that need permissive terms without Vertex-only APIs. It is often fine-tuned for domain tasks on TPU or GPU clusters.
Gemini 3.1 Pro Preview
PreviewLatestGemini 3.1 Pro Preview is Google's preview Gemini 3-series model for complex multimodal reasoning, software engineering behavior, and agentic workflows requiring precise tool use.
Gemini 3.7 Flash
CurrentGemini 3.7 Flash is Google's workhorse model for coding, agents, web development, knowledge work, and multimodal workflows. Google documents availability through the Gemini API, Google AI Studio, Android Studio, Gemini Enterprise Agent Platform, and Gemini Spark, with a 1M-token input context and up to 64K output.
Gemini 2.5 Pro
CurrentGoogle's advanced Gemini model for complex tasks, with official Gemini API documentation calling out deep reasoning and coding capabilities.
Gemini 3.5 Flash
LegacyGemini 3.5 Flash is Google's stable Gemini 3-series Flash model for agentic and coding tasks where teams need strong performance with lower latency and cost than Pro.
Gemini 1.5 Pro
LegacyGoogle DeepMind Gemini 1.5 Pro targets long-context multimodal workloads—large effective context for retrieval-heavy document pipelines, plus image, audio, and video inputs on supported surfaces. It is often paired with Vertex AI or the Gemini API for enterprise workloads on GCP.
Gemini 2.5 Flash-Lite
LegacyGoogle's fastest and most budget-friendly multimodal model in the Gemini 2.5 family, according to the Gemini API model documentation.
Gemini 1.5 Flash
LegacyGemini 1.5 Flash targets low-latency, cost-efficient multimodal chat and retrieval workloads on the Gemini API and Vertex AI. It keeps much of the long-context family behavior with faster responses for interactive apps.
Gemini 1.0 Pro
LegacyGemini 1.0 Pro represents Google’s first broadly marketed Gemini-era general model for text and basic multimodal tasks on Vertex and consumer surfaces. New projects should prefer 1.5+ generations unless constrained by legacy integrations—verify availability.
Gemini 2.0 Flash
LegacyGemini 2.0 Flash is Google’s efficiency-oriented multimodal model generation aimed at fast agentic and interactive experiences. Capabilities and naming evolve—validate against the current Gemini API reference for tool use and context limits.
Recommended
Top current models by information quality score — good defaults when you are not sure where to start.
Anthropic
Claude Opus 5.5
CurrentLatestClaude Opus 5.5 is Anthropic's September 22, 2026 Opus-class model for long-running agentic coding and knowledge work, and the first model in the Claude 5.5 family. The Claude API ID is claude-opus-5-5, with a 1M-token context window, 128K synchronous max output, adaptive thinking that cannot be disabled, and default effort medium. Anthropic prices it at $4 input and $20 output per million tokens and says typical workloads cost about 40% less to run than Opus 5.
OpenAI
GPT-Image-2.5 Sunburst
CurrentLatestGPT-Image-2.5 Sunburst is OpenAI's September 8, 2026 most capable image generation and editing model. The API ID is gpt-image-2.5-sunburst, with dated snapshot gpt-image-2.5-sunburst-2026-09-08. It takes text and image input and outputs images. OpenAI positions it for workflows where editing precision matters most, with quality settings low, medium, high, xhigh, max, and auto. Call it on the Images API or as the model of the Responses API image generation tool.
Alibaba Qwen
Qwen3.8-27B
CurrentLatestQwen3.8-27B is Qwen's deployment-oriented dense multimodal model for coding, professional work, research, and long-horizon agents. The official repository documents 27B parameters, native image and video understanding, flexible reasoning effort, a 262,144-token native context extensible to 1M, and compatibility with Transformers, vLLM, SGLang, and TokenSpeed.
Meta
Muse Glimmer 30B
CurrentLatestMuse Glimmer 30B is Meta Superintelligence Labs' open-weight multimodal model for local agents, coding, tool use, long-horizon reasoning, and image understanding. Meta's model card documents a dense 29.6B-parameter architecture with a dedicated perception encoder, 131,072+ context, text-and-image input, text output, controllable reasoning effort, and Apache 2.0 weights.
OpenAI
GPT-6 Sol
CurrentLatestGPT-6 Sol is OpenAI's September 22, 2026 model for complex coding and agentic workflows, positioned as a faster and cheaper GPT-6 lane than GPT-6 Astra. The API model ID is gpt-6-sol, with a 1,050,000-token context window, 128,000 max output tokens, text and image input, and reasoning.effort from none through max (default medium). Standard text pricing is $2 input and $10 output per million tokens.
Missing a frontier release? Add a model (editors)