GenAIWiki

Structured cards

Model database

Filter by provider, architecture family, or full-text search across descriptions.

Frontier models

Verified flagships and this month’s launches—Claude Opus 5.5, GPT-6 Sol and Luna, GPT-6 Astra, Claude Fable 5.1, Gemini 3.8, and more. Scroll sideways for the full shelf.

OpenAI

GPT-6 Astra

FrontierLatest

GPT-6 Astra is OpenAI's September 2026 flagship for the hardest end-to-end reasoning, coding, computer use, research, and document work. The API model ID is gpt-6-astra, with a 1,050,000-token context window, 128,000 max output tokens, text and image input, and reasoning.effort from low through max. OpenAI's safety overview (September 3, 2026) states Astra is the first broadly deployed OpenAI model to reach the Critical cybersecurity capability level under the Preparedness Framework.

FeaturedUpdated 2 weeks ago
openaigpt-6

Anthropic

Claude Fable 5.1

FrontierLatest

Claude Fable 5.1 is Anthropic's September 1, 2026 generally available flagship for demanding reasoning and long-horizon agentic coding, research, and document work. The API ID is claude-fable-5-1, with a 1M-token context window, 128K max output, adaptive thinking, and default high effort. Anthropic keeps the $10/$50 input/output list price of Fable 5 while cutting cache reads to $0.25 per million tokens.

FeaturedUpdated 2 weeks ago
anthropicclaude

Google

Gemini 3.8 Flash

FrontierLatest

Gemini 3.8 Flash is Google's September 2, 2026 workhorse for software engineering, agentic tasks, and multi-step reasoning at Flash speed. The DeepMind model card documents text, image, audio, and video input, a 1M-token context window, 64K output, and adjustable effort levels. Google keeps the same introductory price as 3.7 Flash: $0.75 input and $3.75 output per million tokens through December 31, 2026.

FeaturedUpdated 2 weeks ago
googlegemini

Google

Gemini 3.8 Live

FrontierLatest

Gemini 3.8 Live is Google's September 15, 2026 default Live API model for low-latency voice agents and real-time dialogue. The stable model code is gemini-3.8-live. It takes text, images, audio, and video, and returns text and audio, with a 131,072-token input limit and 65,536-token output limit. Google positions it for scale and cost efficiency: interleaved reasoning, asynchronous function calling, 97-language mid-conversation switching, and visual grounding while the session keeps talking. Audio sessions are listed at $0.005 per minute input and $0.018 per minute output on the developer Live post.

FeaturedUpdated 6 days ago
googlegemini

TypeSafe AI

Jev

FrontierLatest

Jev is TypeSafe AI's first System One model, launched in early access on September 15, 2026. It is not a chatbot: you send unstructured or structured state plus typed questions and get structured answers your code can branch on. The documented model alias is jev-latest on POST https://api.typesafe.ai/v1/systemone. Question primitives are Choice (up to 255 options), Score, and Noul (yes/no probability). TypeSafe lists $0.042 per million input tokens with output tokens unmetered, and 70–500 ms end-to-end on its launch workloads. Training is Reinforcement Learning for Calibrated Decisions (RLCD).

FeaturedUpdated 7 days ago
typesafejev

OpenAI

GPT-Live-1

FrontierLatest

GPT-Live-1 is OpenAI's full-duplex voice model for real-time spoken conversation. The API ID is gpt-live-1. It listens and speaks at the same time, handles interruptions, and delegates reasoning and tool use to a backend agent instead of doing that work itself. OpenAI's system card (July 8, 2026) introduced it as the paid ChatGPT voice default. The API launch post (September 10, 2026) documents Live sessions, native ASR transcripts and response text, keyword biasing, and turn detection. Voice sessions are billed $0.05 per minute, per second. Backend model and tool usage is billed separately.

FeaturedUpdated 10 days ago
openaigpt-live

Meta

Muse Spark 1.3

FrontierLatest

Muse Spark 1.3 is Meta's September 2, 2026 hosted model for long-horizon agentic and coding work. The API ID is muse-spark-1.3, with a 1,048,576-token context window, text/image/video/PDF input, and text output. Meta documents it in Muse Code and Meta Model API, with reasoning effort including max on the Standard tier. Official docs still list muse-spark-1.2 as the Muse Code default while recommending 1.3 for new API work.

FeaturedUpdated 2 weeks ago
metamuse

DeepSeek

DeepSeek-V4.1-Flash

FrontierLatest

DeepSeek-V4.1-Flash is DeepSeek's September 10, 2026 multimodal Mixture-of-Experts model and the smallest member of its new architecture family. The live API ID is deepseek-flash. Official materials describe a 552B-parameter Causal Encoder–Decoder MoE that activates 8B parameters on input and 16B on output, native image-and-text understanding, and a 1-million-token context. DeepSeek says V4.1-Flash outperforms DeepSeek-V4-Pro on performance, cost, speed, and total runtime. Weights are on Hugging Face under MIT.

FeaturedUpdated 10 days ago
deepseekv4.1

SpaceXAI

Grok 4.6

FrontierLatest

Grok 4.6 is SpaceXAI's frontier model for coding, long-running agents, knowledge work, and interactive visual projects. Official documentation lists text and image input, text output, a 500,000-token context window, configurable low through xhigh reasoning, function calling, web and X search, code execution, and the API model ID grok-4.6.

FeaturedUpdated 5 weeks ago
xaispacexai

Z.ai

GLM-5.3-Flash

FrontierLatest

GLM-5.3-Flash is Z.ai's first natively multimodal GLM-5 model: a 320B-parameter mixture-of-experts checkpoint with 18B active parameters, hybrid sparse and linear attention, and MIT-licensed weights. Z.ai documents API access as glm-5.3-flash, Coding Plan and ZCode availability, and Hugging Face weights for SGLang, vLLM, and TokenSpeed.

FeaturedUpdated 3 weeks ago
zaiglm

Alibaba Qwen

Qwen3.8-Max

FrontierLatest

Qwen3.8-Max is Alibaba Qwen's hosted 2.4-trillion-parameter mixture-of-experts flagship for coding, professional work, multimodal analysis, and long-horizon agents. Effective September 5, 2026 (UTC+8), the qwen3.8-max endpoint automatically serves the September 2 snapshot qwen3.8-max-0902 (alias qwen3.8-max-2026-09-02). QwenCloud says 0902 improves coding depth, multi-tool agent delivery, and visual understanding while keeping the 1M context, thinking mode, tools, and unchanged billing.

FeaturedUpdated 10 days ago
qwenalibaba

Microsoft AI

MAI-Image-2.6

FrontierLatest

MAI-Image-2.6 is Microsoft AI's September 4, 2026 highest-precision image generation and editing model, in public preview on Microsoft Foundry and MAI Playground. Microsoft says it ranks No. 2 for text-to-image and image editing on Arena and Artificial Analysis as of September 4, with multi-image reference editing, web grounding, dynamic aspect ratios, and up to 1.5K resolution.

FeaturedUpdated 2 weeks ago
microsoftmai

By company

Latest verified models from each frontier company, summarized on one page.

All companies

All models

Filter and paginate the full catalog. Tabs control lifecycle scope.

Google

Gemini 3.8 Live

CurrentLatest

Gemini 3.8 Live is Google's September 15, 2026 default Live API model for low-latency voice agents and real-time dialogue. The stable model code is gemini-3.8-live. It takes text, images, audio, and video, and returns text and audio, with a 131,072-token input limit and 65,536-token output limit. Google positions it for scale and cost efficiency: interleaved reasoning, asynchronous function calling, 97-language mid-conversation switching, and visual grounding while the session keeps talking. Audio sessions are listed at $0.005 per minute input and $0.018 per minute output on the developer Live post.

FeaturedUpdated 6 days ago
googlegemini

Google

Gemini 3.8 Live Extended Thinking

CurrentLatest

Gemini 3.8 Live Extended Thinking is Google's September 15, 2026 high-reasoning Live API model for complex multi-step work during a live voice session. The stable model code is gemini-3.8-live-extended-thinking. Same 131,072 / 65,536 token limits and text/image/audio/video in, text and audio out as Gemini 3.8 Live. It runs background reasoning and asynchronous tool calls while streaming audio. Google cites #1 on Artificial Analysis Speech-to-Speech Quality Index (82.6), 68.6% on τ-Voice, 35.1% on Sierra τ-Voice-banking, and 97.7% on Big Bench Audio. Same listed audio session rates as 3.8 Live.

FeaturedUpdated 7 days ago
googlegemini

Google

Gemini 3.8 Flash

CurrentLatest

Gemini 3.8 Flash is Google's September 2, 2026 workhorse for software engineering, agentic tasks, and multi-step reasoning at Flash speed. The DeepMind model card documents text, image, audio, and video input, a 1M-token context window, 64K output, and adjustable effort levels. Google keeps the same introductory price as 3.7 Flash: $0.75 input and $3.75 output per million tokens through December 31, 2026.

FeaturedUpdated 2 weeks ago
googlegemini

Google

Gemini 3.5 Transcribe

CurrentLatest

Gemini 3.5 Transcribe is Google's speech-to-text model for precise, formatted transcription of live and recorded audio. Google documents public-preview access in the Gemini API, Google AI Studio, Google Antigravity, and Gemini Enterprise Agent Platform, with product surfaces including Rambler on Android and the Gemini app on macOS.

FeaturedUpdated 3 weeks ago
googlegemini

Google

Gemini Omni 1.1 Flash

CurrentLatest

Gemini Omni 1.1 Flash is Google's production video generation and editing model for developers. The August 27, 2026 announcement documents scene extension from up to 10 seconds of prior context to a 40-second cumulative clip, first-and-last-frame interpolation, 360p drafts, upscaling to 1080p or 4K, and video references, with the API model ID gemini-omni-1.1-flash.

FeaturedUpdated 3 weeks ago
googlegemini

Google

WeatherNext 3

CurrentLatest

WeatherNext 3 is Google DeepMind and Google Research's September 3, 2026 global weather AI model. It ingests live geostationary satellite mosaics and station observations to produce hourly forecasts, with surface variables down to 5 km and atmospheric variables at 25 km. Google is integrating it into Search, the Gemini app, Maps, the Maps Platform Weather API, Earth Engine, BigQuery, and Cloud Storage.

FeaturedUpdated 2 weeks ago
googledeepmind

Google

Gemini 3.8 Flash Cyber

CurrentLatest

Gemini 3.8 Flash Cyber is Google's September 2, 2026 defensive-security variant of Gemini 3.8 Flash, gated through the Fairwind Program for trusted government authorities, critical-infrastructure operators, and software maintainers. Google positions it for autonomous vulnerability discovery and automated patching, not public Gemini API self-serve.

FeaturedUpdated 2 weeks ago
googlegemini

Google

TimesFM-3

CurrentLatest

TimesFM-3 is Google Research's 330-million-parameter time-series foundation model for zero-shot univariate and multivariate forecasting. The August 31, 2026 announcement documents native multivariate forecasting in a single forward pass, past and past-future covariates, a pretraining corpus of more than 1 trillion time points, and availability on GitHub and Hugging Face, with BigQuery integration planned in the coming weeks.

FeaturedUpdated 3 weeks ago
googlegoogle-research

Google

Gemini 3.1 Flash-Lite

CurrentLatest

Gemini 3.1 Flash-Lite is Google's stable Gemini 3-series workhorse model for cost-efficient, high-volume multimodal workloads.

Updated 5 weeks ago
googleflash-lite

Google

Gemma 2 27B

CurrentLatest

Gemma 2 27B is Google’s open-weights Gemma family checkpoint balancing quality and deployability for research and product teams that need permissive terms without Vertex-only APIs. It is often fine-tuned for domain tasks on TPU or GPU clusters.

Updated 5 weeks ago
open-weightsgoogle

Google

Gemini 3.7 Flash

Current

Gemini 3.7 Flash is Google's workhorse model for coding, agents, web development, knowledge work, and multimodal workflows. Google documents availability through the Gemini API, Google AI Studio, Android Studio, Gemini Enterprise Agent Platform, and Gemini Spark, with a 1M-token input context and up to 64K output.

FeaturedUpdated 2 weeks ago
googlegemini

Google

Gemini 2.5 Pro

Current

Google's advanced Gemini model for complex tasks, with official Gemini API documentation calling out deep reasoning and coding capabilities.

FeaturedUpdated 5 weeks ago
frontiergoogle

Recommended

Top current models by information quality score — good defaults when you are not sure where to start.

Anthropic

Claude Opus 5.5

CurrentLatest

Claude Opus 5.5 is Anthropic's September 22, 2026 Opus-class model for long-running agentic coding and knowledge work, and the first model in the Claude 5.5 family. The Claude API ID is claude-opus-5-5, with a 1M-token context window, 128K synchronous max output, adaptive thinking that cannot be disabled, and default effort medium. Anthropic prices it at $4 input and $20 output per million tokens and says typical workloads cost about 40% less to run than Opus 5.

FeaturedUpdated 1 day ago
anthropicclaude

OpenAI

GPT-Image-2.5 Sunburst

CurrentLatest

GPT-Image-2.5 Sunburst is OpenAI's September 8, 2026 most capable image generation and editing model. The API ID is gpt-image-2.5-sunburst, with dated snapshot gpt-image-2.5-sunburst-2026-09-08. It takes text and image input and outputs images. OpenAI positions it for workflows where editing precision matters most, with quality settings low, medium, high, xhigh, max, and auto. Call it on the Images API or as the model of the Responses API image generation tool.

FeaturedUpdated 2 weeks ago
openaigpt-image

Alibaba Qwen

Qwen3.8-27B

CurrentLatest

Qwen3.8-27B is Qwen's deployment-oriented dense multimodal model for coding, professional work, research, and long-horizon agents. The official repository documents 27B parameters, native image and video understanding, flexible reasoning effort, a 262,144-token native context extensible to 1M, and compatibility with Transformers, vLLM, SGLang, and TokenSpeed.

FeaturedUpdated 5 weeks ago
alibabaqwen

Meta

Muse Glimmer 30B

CurrentLatest

Muse Glimmer 30B is Meta Superintelligence Labs' open-weight multimodal model for local agents, coding, tool use, long-horizon reasoning, and image understanding. Meta's model card documents a dense 29.6B-parameter architecture with a dedicated perception encoder, 131,072+ context, text-and-image input, text output, controllable reasoning effort, and Apache 2.0 weights.

FeaturedUpdated 5 weeks ago
metamuse

OpenAI

GPT-6 Sol

CurrentLatest

GPT-6 Sol is OpenAI's September 22, 2026 model for complex coding and agentic workflows, positioned as a faster and cheaper GPT-6 lane than GPT-6 Astra. The API model ID is gpt-6-sol, with a 1,050,000-token context window, 128,000 max output tokens, text and image input, and reasoning.effort from none through max (default medium). Standard text pricing is $2 input and $10 output per million tokens.

FeaturedUpdated 1 day ago
openaigpt-6

Missing a frontier release? Add a model (editors)