Structured cards
Model database
Filter by provider, architecture family, or full-text search across descriptions.
Frontier models
Verified flagships and recent launches across major providers—scroll sideways for the full shelf.
All models
Filter and paginate the full catalog. Tabs control lifecycle scope.
AWS
Amazon Titan Text Premier
CurrentLatestTitan Text Premier is AWS’s managed text model for Bedrock workloads emphasizing integration with guardrails, knowledge bases, and private data patterns. It targets enterprise RAG and internal assistants rather than frontier creative writing.
Meta
Llama 3.1 8B Instruct
CurrentLatestLlama 3.1 8B Instruct is a small open-weights model for edge laptops, single-GPU servers, and ultra-low-latency assistants. Quality per dollar is competitive for simple tasks but not for frontier reasoning.
Gemma 2 27B
CurrentLatestGemma 2 27B is Google’s open-weights Gemma family checkpoint balancing quality and deployability for research and product teams that need permissive terms without Vertex-only APIs. It is often fine-tuned for domain tasks on TPU or GPU clusters.
Mistral AI
Mixtral 8x22B
CurrentLatestCatalog entry for this named release; see the provider’s official documentation for modalities, pricing, and context limits.
Community
LLaVA
CurrentLatestCatalog entry for this named release; see the provider’s official documentation for modalities, pricing, and context limits.
xAI
Grok 1.5
CurrentLatestCatalog entry for this named release; see the provider’s official documentation for modalities, pricing, and context limits.
DeepSeek
DeepSeek Coder V2
CurrentLatestCatalog entry for this named release; see the provider’s official documentation for modalities, pricing, and context limits.
Databricks
DBRX
CurrentLatestCatalog entry for this named release; see the provider’s official documentation for modalities, pricing, and context limits.
Microsoft AI
MAI-Image-2.5-Pro
PreviewLatestMAI-Image-2.5-Pro is Microsoft AI's highest-fidelity image model to date, released in public preview on July 23, 2026. Microsoft positions it for hero imagery, detailed natural-language editing, and precise in-image text when quality matters more than the fastest or cheapest route.
Microsoft AI
MAI-Voice-2-Flash
PreviewLatestMAI-Voice-2-Flash is Microsoft AI's public-preview speech model for high-volume, low-latency voice experiences. Microsoft reports that it is twice as fast and 32 percent cheaper than MAI-Voice-2 while retaining natural prosody and high acoustic quality.
Gemini 3.1 Pro Preview
PreviewLatestGemini 3.1 Pro Preview is Google's preview Gemini 3-series model for complex multimodal reasoning, software engineering behavior, and agentic workflows requiring precise tool use.
Microsoft AI
MAI-Thinking-1
PreviewLatestMicrosoft AI's frontier reasoning model in the MAI family, announced for difficult prompts, science, math, and complex planning workloads, with Microsoft Foundry access documented as private preview.
Anthropic
Claude Mythos 5
PreviewLatestAnthropic's experimental Claude model for open-ended scientific inquiry and long-horizon reasoning, documented as limited to approved Project Glasswing researchers.
Gemini 2.5 Pro
CurrentGoogle's advanced Gemini model for complex tasks, with official Gemini API documentation calling out deep reasoning and coding capabilities.
xAI
Grok 4.3
CurrentxAI's documented default for general chat workloads, described in xAI docs as the most intelligent and fastest Grok model for non-specialized use cases.
Anthropic
Claude Mythos Preview
PreviewCatalog entry for this named release; see the provider’s official documentation for modalities, pricing, and context limits.
Alibaba Qwen
Qwen3.6-27B
LegacyQwen3.6-27B is Alibaba Qwen's Apache 2.0 open-weight multimodal model for coding, repository-level reasoning, tool-driven workflows, and long-context tasks. The official model card documents a 27B language model with a vision encoder, 262,144 tokens of native context, optional extension to 1,010,000 tokens, and support in Transformers, vLLM, SGLang, and KTransformers.
Z.ai
GLM-5.2
LegacyGLM-5.2 is Z.ai's open-weight model for long-horizon reasoning, coding, and agent workflows. Official Z.ai and Hugging Face materials document a one-million-token context window, flexible reasoning effort, a mixture-of-experts architecture, and MIT-licensed weights.
xAI
Grok 4.5
LegacyGrok 4.5 is xAI's frontier model for coding, agentic tasks, and knowledge work, available through the xAI API and in Cursor as a first-party model option.
Alibaba Qwen
Qwen3.8-Max-Preview
LegacyQwen3.8-Max-Preview is Alibaba's July 2026 preview of its newest Qwen Max model for coding, agentic, and general reasoning workflows. Alibaba's Qwen Code materials identify qwen3.8-max as qwen3.8-max-preview; because this is a preview, endpoint behavior, limits, pricing, and availability should be treated as changeable.
Anthropic
Claude Opus 4.7
LegacyAnthropic's most capable generally available Claude model for complex reasoning and agentic coding, documented in the Claude model overview.
Anthropic
Claude Opus 4.8
LegacyAnthropic's current Opus-tier Claude model, documented for complex reasoning, coding, and multimodal enterprise workloads below the newer Fable tier.
OpenAI
GPT-5.5
LegacyOpenAI's current flagship model for complex reasoning, coding, and professional work, documented in the OpenAI API model guide as the default starting point for high-complexity workloads.
Gemini 3.5 Flash
LegacyGemini 3.5 Flash is Google's stable Gemini 3-series Flash model for agentic and coding tasks where teams need strong performance with lower latency and cost than Pro.
Recommended
Top current models by information quality score — good defaults when you are not sure where to start.
Sarvam AI
Sarvam 30B
CurrentLatestSarvam 30B is a 30B parameter Mixture-of-Experts chat and reasoning model from Sarvam AI, optimized for Indian languages, real-time conversation, high-throughput voice-agent pipelines, coding, and practical deployment. Sarvam documents 2.4B active parameters per token, 16T tokens of pre-training data, a 64K context window, Grouped Query Attention, Apache 2.0 open weights, and OpenAI-compatible chat completions.
MiniMax
MiniMax M3
CurrentLatestMiniMax M3 is a June 2026 open-weight multimodal model for coding, agentic workflows, computer use, and long-context work. MiniMax documents a one-million-token context window, native image and video understanding, and deployment through hosted or downloadable model paths.
Alibaba Qwen
Qwen3.8-2.4T-A95B
CurrentLatestQwen3.8-2.4T-A95B is Qwen's downloadable 2.4-trillion-parameter sparse mixture-of-experts model with 95B activated parameters. The official model card documents text-only input and output, mandatory thinking, 262,144 native context extensible to roughly 1.01M, configurable reasoning effort, and a dedicated Qwen3.8-Max license.
Missing a frontier release? Add a model (editors)