GenAIWiki

Structured cards

Model database

Filter by provider, architecture family, or full-text search across descriptions.

Frontier models

Verified flagships and this month’s launches—Claude Sonnet 5.5, GPT-6.1 Sol, GPT-6 Luna and Astra, Claude Opus 5.5, Gemini 3.8, and more. Scroll sideways for the full shelf.

Anthropic

Claude Sonnet 5.5

FrontierLatest

Claude Sonnet 5.5 is Anthropic's September 28, 2026 Sonnet model, documented as a clear upgrade over Claude Sonnet 5 that runs more than 30% faster and costs up to 30% less for most work. The Claude API ID is claude-sonnet-5-5. It has a 1M-token context window, 128K synchronous max output, text and image input, text output, adaptive thinking, and default effort high. Standard price is $2 input and $10 output per million tokens.

FeaturedUpdated 4 days ago
anthropicsonnet

OpenAI

GPT-6.1 Sol

FrontierLatest

GPT-6.1 Sol is OpenAI's current Sol model, positioned as near-Astra performance for complex work at a lower cost than Astra. The API ID is gpt-6.1-sol. It accepts text and image input, returns text, and has a 1,050,000-token context window, 128,000 max output, and a knowledge cutoff of April 30, 2026. Standard price is $2 input, $0.10 cached input, $2.50 cache writes, and $10 output per million tokens.

FeaturedUpdated 4 days ago
openaigpt-6

Google

Gemini 3.8 Flash TTS

FrontierLatest

Gemini 3.8 Flash TTS is Google's September 2026 creative and studio text-to-speech model. The model code is gemini-3.8-flash-tts. It takes text input and returns audio, with an 8,192-token input limit, a 16,384-token output limit, and 130 languages. Default unary audio is WAV. Google's Gemini API pricing page does not list a paid rate for this model ID.

FeaturedUpdated 4 days ago
googlegemini

ElevenLabs

Eleven v3

FrontierLatest

Eleven v3 is ElevenLabs' documented flagship for human-like, expressive speech. The model ID is eleven_v3. It supports 70+ languages and is used with the Text to Speech API and the Text to Dialogue API. The official models page lists a 5,000-character request limit, about five minutes of audio. Help Center still calls Eleven v3 the latest speech model; Eleven v4 is not on the models page.

FeaturedUpdated 4 days ago
elevenlabstts

Sarvam AI

Sarvam Vision

FrontierLatest

Sarvam Vision is Sarvam AI's document vision model for Indic and English pages. The documented API model ID is sarvam-vision, a 3B vision-language model covering 23 languages (22 Indian languages plus English). Document AI digitise and extract accept up to 10 pages per PDF and 200 MB per file, at 10 requests per minute. Sarvam's public pricing page lists ₹0.50 per page. The September 24, 2026 post describes Vision 2.1 as a capability update; the docs still use sarvam-vision.

FeaturedUpdated 4 days ago
sarvamvision

OpenAI

GPT-6 Astra

FrontierLatest

GPT-6 Astra is OpenAI's September 2026 flagship for the hardest end-to-end reasoning, coding, computer use, research, and document work. The API model ID is gpt-6-astra, with a 1,050,000-token context window, 128,000 max output tokens, text and image input, and reasoning.effort from low through max. OpenAI's safety overview (September 3, 2026) states Astra is the first broadly deployed OpenAI model to reach the Critical cybersecurity capability level under the Preparedness Framework.

FeaturedUpdated 4 weeks ago
openaigpt-6

Anthropic

Claude Fable 5.1

FrontierLatest

Claude Fable 5.1 is Anthropic's September 1, 2026 generally available flagship for demanding reasoning and long-horizon agentic coding, research, and document work. The API ID is claude-fable-5-1, with a 1M-token context window, 128K max output, adaptive thinking, and default high effort. Anthropic keeps the $10/$50 input/output list price of Fable 5 while cutting cache reads to $0.25 per million tokens.

FeaturedUpdated 4 weeks ago
anthropicclaude

Google

Gemini 3.8 Flash

FrontierLatest

Gemini 3.8 Flash is Google's September 2, 2026 workhorse for software engineering, agentic tasks, and multi-step reasoning at Flash speed. The DeepMind model card documents text, image, audio, and video input, a 1M-token context window, 64K output, and adjustable effort levels. Google keeps the same introductory price as 3.7 Flash: $0.75 input and $3.75 output per million tokens through December 31, 2026.

FeaturedUpdated 4 weeks ago
googlegemini

Google

Gemini 3.8 Live

FrontierLatest

Gemini 3.8 Live is Google's September 15, 2026 default Live API model for low-latency voice agents and real-time dialogue. The stable model code is gemini-3.8-live. It takes text, images, audio, and video, and returns text and audio, with a 131,072-token input limit and 65,536-token output limit. Google positions it for scale and cost efficiency: interleaved reasoning, asynchronous function calling, 97-language mid-conversation switching, and visual grounding while the session keeps talking. Audio sessions are listed at $0.005 per minute input and $0.018 per minute output on the developer Live post.

FeaturedUpdated 2 weeks ago
googlegemini

TypeSafe AI

Jev

FrontierLatest

Jev is TypeSafe AI's first System One model, launched in early access on September 15, 2026. It is not a chatbot: you send unstructured or structured state plus typed questions and get structured answers your code can branch on. The documented model alias is jev-latest on POST https://api.typesafe.ai/v1/systemone. Question primitives are Choice (up to 255 options), Score, and Noul (yes/no probability). TypeSafe lists $0.042 per million input tokens with output tokens unmetered, and 70–500 ms end-to-end on its launch workloads. Training is Reinforcement Learning for Calibrated Decisions (RLCD).

FeaturedUpdated 3 weeks ago
typesafejev

OpenAI

GPT-Live-1

FrontierLatest

GPT-Live-1 is OpenAI's full-duplex voice model for real-time spoken conversation. The API ID is gpt-live-1. It listens and speaks at the same time, handles interruptions, and delegates reasoning and tool use to a backend agent instead of doing that work itself. OpenAI's system card (July 8, 2026) introduced it as the paid ChatGPT voice default. The API launch post (September 10, 2026) documents Live sessions, native ASR transcripts and response text, keyword biasing, and turn detection. Voice sessions are billed $0.05 per minute, per second. Backend model and tool usage is billed separately.

FeaturedUpdated 3 weeks ago
openaigpt-live

Meta

Muse Spark 1.3

FrontierLatest

Muse Spark 1.3 is Meta's September 2, 2026 hosted model for long-horizon agentic and coding work. The API ID is muse-spark-1.3, with a 1,048,576-token context window, text/image/video/PDF input, and text output. Meta documents it in Muse Code and Meta Model API, with reasoning effort including max on the Standard tier. Official docs still list muse-spark-1.2 as the Muse Code default while recommending 1.3 for new API work.

FeaturedUpdated 4 weeks ago
metamuse

DeepSeek

DeepSeek-V4.1-Flash

FrontierLatest

DeepSeek-V4.1-Flash is DeepSeek's September 10, 2026 multimodal Mixture-of-Experts model and the smallest member of its new architecture family. The live API ID is deepseek-flash. Official materials describe a 552B-parameter Causal Encoder–Decoder MoE that activates 8B parameters on input and 16B on output, native image-and-text understanding, and a 1-million-token context. DeepSeek says V4.1-Flash outperforms DeepSeek-V4-Pro on performance, cost, speed, and total runtime. Weights are on Hugging Face under MIT.

FeaturedUpdated 3 weeks ago
deepseekv4.1

SpaceXAI

Grok 4.6

FrontierLatest

Grok 4.6 is SpaceXAI's frontier model for coding, long-running agents, knowledge work, and interactive visual projects. Official documentation lists text and image input, text output, a 500,000-token context window, configurable low through xhigh reasoning, function calling, web and X search, code execution, and the API model ID grok-4.6.

FeaturedUpdated 6 weeks ago
xaispacexai

Z.ai

GLM-5.3-Flash

FrontierLatest

GLM-5.3-Flash is Z.ai's first natively multimodal GLM-5 model: a 320B-parameter mixture-of-experts checkpoint with 18B active parameters, hybrid sparse and linear attention, and MIT-licensed weights. Z.ai documents API access as glm-5.3-flash, Coding Plan and ZCode availability, and Hugging Face weights for SGLang, vLLM, and TokenSpeed.

FeaturedUpdated 5 weeks ago
zaiglm

Alibaba Qwen

Qwen3.8-Max

FrontierLatest

Qwen3.8-Max is Alibaba Qwen's hosted 2.4-trillion-parameter mixture-of-experts flagship for coding, professional work, multimodal analysis, and long-horizon agents. Effective September 5, 2026 (UTC+8), the qwen3.8-max endpoint automatically serves the September 2 snapshot qwen3.8-max-0902 (alias qwen3.8-max-2026-09-02). QwenCloud says 0902 improves coding depth, multi-tool agent delivery, and visual understanding while keeping the 1M context, thinking mode, tools, and unchanged billing.

FeaturedUpdated 3 weeks ago
qwenalibaba

Microsoft AI

MAI-Image-2.6

FrontierLatest

MAI-Image-2.6 is Microsoft AI's September 4, 2026 highest-precision image generation and editing model, in public preview on Microsoft Foundry and MAI Playground. Microsoft says it ranks No. 2 for text-to-image and image editing on Arena and Artificial Analysis as of September 4, with multi-image reference editing, web grounding, dynamic aspect ratios, and up to 1.5K resolution.

FeaturedUpdated 4 weeks ago
microsoftmai

By company

Latest verified models from each frontier company, summarized on one page.

All companies

All models

Filter and paginate the full catalog. Tabs control lifecycle scope.

ElevenLabs

Eleven v3

CurrentLatest

Eleven v3 is ElevenLabs' documented flagship for human-like, expressive speech. The model ID is eleven_v3. It supports 70+ languages and is used with the Text to Speech API and the Text to Dialogue API. The official models page lists a 5,000-character request limit, about five minutes of audio. Help Center still calls Eleven v3 the latest speech model; Eleven v4 is not on the models page.

FeaturedUpdated 4 days ago
elevenlabstts

ElevenLabs

Eleven v3 Conversational

CurrentLatest

Eleven v3 Conversational is ElevenLabs' expressive realtime speech model. The model ID is eleven_v3_conversational. ElevenLabs documents about 280ms model latency, excluding application and network latency, and 70+ languages. It is the model the docs recommend for support agents, assistants, and interactive characters, including the Text to Dialogue WebSocket.

Updated 4 days ago
elevenlabstts

ElevenLabs

Eleven Flash v2.5

CurrentLatest

Eleven Flash v2.5 is ElevenLabs' ultra-fast speech model for realtime use. The model ID is eleven_flash_v2_5. ElevenLabs documents about 75ms model latency and support for the Multilingual v2 languages plus Hungarian, Norwegian, and Vietnamese. The models page lists a 40,000-character request limit, about 40 minutes. Number normalization is off by default to keep latency low.

Updated 4 days ago
elevenlabstts

ElevenLabs

Scribe v2

CurrentLatest

Scribe v2 is ElevenLabs' documented state-of-the-art speech recognition model. The model ID is scribe_v2. It supports 90+ languages. The realtime sibling on the same models page is scribe_v2_realtime. Scribe v1 is listed as outclassed by the v2 models.

Updated 4 days ago
elevenlabsstt

ElevenLabs

Eleven Music v2

CurrentLatest

Eleven Music v2 is ElevenLabs' current music model. The model ID is music_v2. It generates studio-grade music from text prompts, composition plans, and previously generated songs, in English, Spanish, German, Japanese, and more. ElevenLabs says music_v1 is outclassed by music_v2.

Updated 4 days ago
elevenlabsmusic

Recommended

Top current models by information quality score — good defaults when you are not sure where to start.

Anthropic

Claude Opus 5.5

CurrentLatest

Claude Opus 5.5 is Anthropic's September 22, 2026 Opus-class model for long-running agentic coding and knowledge work, and the first model in the Claude 5.5 family. The Claude API ID is claude-opus-5-5, with a 1M-token context window, 128K synchronous max output, adaptive thinking that cannot be disabled, and default effort medium. Anthropic prices it at $4 input and $20 output per million tokens and says typical workloads cost about 40% less to run than Opus 5.

FeaturedUpdated 12 days ago
anthropicclaude

OpenAI

GPT-Image-2.5 Sunburst

CurrentLatest

GPT-Image-2.5 Sunburst is OpenAI's September 8, 2026 most capable image generation and editing model. The API ID is gpt-image-2.5-sunburst, with dated snapshot gpt-image-2.5-sunburst-2026-09-08. It takes text and image input and outputs images. OpenAI positions it for workflows where editing precision matters most, with quality settings low, medium, high, xhigh, max, and auto. Call it on the Images API or as the model of the Responses API image generation tool.

FeaturedUpdated 4 weeks ago
openaigpt-image

Alibaba Qwen

Qwen3.8-27B

CurrentLatest

Qwen3.8-27B is Qwen's deployment-oriented dense multimodal model for coding, professional work, research, and long-horizon agents. The official repository documents 27B parameters, native image and video understanding, flexible reasoning effort, a 262,144-token native context extensible to 1M, and compatibility with Transformers, vLLM, SGLang, and TokenSpeed.

FeaturedUpdated 6 weeks ago
alibabaqwen

Meta

Muse Glimmer 30B

CurrentLatest

Muse Glimmer 30B is Meta Superintelligence Labs' open-weight multimodal model for local agents, coding, tool use, long-horizon reasoning, and image understanding. Meta's model card documents a dense 29.6B-parameter architecture with a dedicated perception encoder, 131,072+ context, text-and-image input, text output, controllable reasoning effort, and Apache 2.0 weights.

FeaturedUpdated 7 weeks ago
metamuse

Missing a frontier release? Add a model (editors)