Structured cards
Model database
Filter by provider, architecture family, or full-text search across descriptions.
Frontier models
Verified flagships and recent launches across major providers—scroll sideways for the full shelf.
All models
Filter and paginate the full catalog. Tabs control lifecycle scope.
Alibaba Qwen
Qwen3.6-27B
LegacyQwen3.6-27B is Alibaba Qwen's Apache 2.0 open-weight multimodal model for coding, repository-level reasoning, tool-driven workflows, and long-context tasks. The official model card documents a 27B language model with a vision encoder, 262,144 tokens of native context, optional extension to 1,010,000 tokens, and support in Transformers, vLLM, SGLang, and KTransformers.
Z.ai
GLM-5.2
LegacyGLM-5.2 is Z.ai's open-weight model for long-horizon reasoning, coding, and agent workflows. Official Z.ai and Hugging Face materials document a one-million-token context window, flexible reasoning effort, a mixture-of-experts architecture, and MIT-licensed weights.
Alibaba Qwen
Qwen3.8-Max-Preview
LegacyQwen3.8-Max-Preview is Alibaba's July 2026 preview of its newest Qwen Max model for coding, agentic, and general reasoning workflows. Alibaba's Qwen Code materials identify qwen3.8-max as qwen3.8-max-preview; because this is a preview, endpoint behavior, limits, pricing, and availability should be treated as changeable.
xAI
Grok 4.5
LegacyGrok 4.5 is xAI's frontier model for coding, agentic tasks, and knowledge work, available through the xAI API and in Cursor as a first-party model option.
Anthropic
Claude Opus 4.7
LegacyAnthropic's most capable generally available Claude model for complex reasoning and agentic coding, documented in the Claude model overview.
Anthropic
Claude Opus 4.8
LegacyAnthropic's current Opus-tier Claude model, documented for complex reasoning, coding, and multimodal enterprise workloads below the newer Fable tier.
OpenAI
GPT-5.5
LegacyOpenAI's current flagship model for complex reasoning, coding, and professional work, documented in the OpenAI API model guide as the default starting point for high-complexity workloads.
Gemini 3.5 Flash
LegacyGemini 3.5 Flash is Google's stable Gemini 3-series Flash model for agentic and coding tasks where teams need strong performance with lower latency and cost than Pro.
Gemini 1.5 Pro
LegacyGoogle DeepMind Gemini 1.5 Pro targets long-context multimodal workloads—large effective context for retrieval-heavy document pipelines, plus image, audio, and video inputs on supported surfaces. It is often paired with Vertex AI or the Gemini API for enterprise workloads on GCP.
DeepSeek
DeepSeek-V3
LegacyDeepSeek-V3 is a large-scale language model family noted for strong coding and math performance under open or research-friendly terms (verify the exact license for your deployment). Teams adopt it for cost-sensitive research, self-hosted inference, or comparison against frontier APIs.
Anthropic
Claude 3.5 Sonnet
LegacyAnthropic’s balanced Sonnet-tier model tuned for long-context reasoning, careful instruction following, and strong performance on coding and analysis workloads. It is a common enterprise choice on the Anthropic API and on AWS Bedrock when teams need large context for RAG and document review.
Mistral AI
Mistral Large 2
LegacyMistral’s frontier-class multilingual model emphasizing JSON adherence, agent-friendly behavior, and competitive reasoning within the Mistral API ecosystem. European teams often evaluate it for GDPR-adjacent deployment patterns alongside US-hosted alternatives.
OpenAI
GPT-5.4
LegacyOpenAI's GPT-5.4 model, documented in the official OpenAI API model guide as part of the current GPT-5 family below the GPT-5.5 flagship lane.
OpenAI
GPT-4o
LegacyGPT-4o is an OpenAI multimodal model that accepts text and image inputs and produces text. It supports streaming, function calling, Structured Outputs, fine-tuning, and predicted outputs for vision-heavy assistants and structured extraction workflows.
xAI
Grok-2
LegacyGrok-2 is xAI’s flagship chat model positioned for real-time knowledge integrations and high-throughput conversational products on xAI’s API. Availability and pricing evolve—treat capabilities as vendor-specific.
Alibaba
Qwen 2.5 72B Instruct
LegacyQwen 2.5 72B Instruct is a large multilingual open-weights model from Alibaba’s Qwen family with strong coding and general chat performance. Common in APAC deployments and on Hugging Face inference endpoints—check license terms for commercial use.
Anthropic
Claude Sonnet 4.6
LegacyAnthropic's Sonnet-tier model documented as the best combination of speed and intelligence in the Claude model overview.
OpenAI
o1
LegacyOpenAI’s o1 series emphasizes extended internal reasoning before answering—useful for competition-style math, complex debugging, and multi-step planning where latency is acceptable. It behaves differently from standard chat models: tune prompts for chain-of-thought style tasks and measure time-to-first-token.
Gemini 2.5 Flash-Lite
LegacyGoogle's fastest and most budget-friendly multimodal model in the Gemini 2.5 family, according to the Gemini API model documentation.
DeepSeek
DeepSeek-V3.2
LegacyDeepSeek's documented successor to the V3.2 experimental line, positioned in official DeepSeek API news as live on app, web, and API.
OpenAI
GPT-5.4 mini
LegacyOpenAI's smaller GPT-5.4 mini model, documented in the official OpenAI API model guide for lower-latency or lower-cost GPT-5 family workloads.
OpenAI
GPT-5.4 nano
LegacyOpenAI's smallest GPT-5.4 nano model, documented in the official OpenAI API model guide for very low-latency or economical GPT-5 family routing.
Gemini 1.5 Flash
LegacyGemini 1.5 Flash targets low-latency, cost-efficient multimodal chat and retrieval workloads on the Gemini API and Vertex AI. It keeps much of the long-context family behavior with faster responses for interactive apps.
Anthropic
Claude 3.5 Haiku
LegacyClaude 3.5 Haiku is Anthropic’s fast, cost-efficient tier for high-volume classification, routing, and simple chat. It targets latency-sensitive paths and agent pre-processing before escalating to Sonnet-class models.
Recommended
Top current models by information quality score — good defaults when you are not sure where to start.
Sarvam AI
Sarvam 30B
CurrentLatestSarvam 30B is a 30B parameter Mixture-of-Experts chat and reasoning model from Sarvam AI, optimized for Indian languages, real-time conversation, high-throughput voice-agent pipelines, coding, and practical deployment. Sarvam documents 2.4B active parameters per token, 16T tokens of pre-training data, a 64K context window, Grouped Query Attention, Apache 2.0 open weights, and OpenAI-compatible chat completions.
MiniMax
MiniMax M3
CurrentLatestMiniMax M3 is a June 2026 open-weight multimodal model for coding, agentic workflows, computer use, and long-context work. MiniMax documents a one-million-token context window, native image and video understanding, and deployment through hosted or downloadable model paths.
Alibaba Qwen
Qwen3.8-2.4T-A95B
CurrentLatestQwen3.8-2.4T-A95B is Qwen's downloadable 2.4-trillion-parameter sparse mixture-of-experts model with 95B activated parameters. The official model card documents text-only input and output, mandatory thinking, 262,144 native context extensible to roughly 1.01M, configurable reasoning effort, and a dedicated Qwen3.8-Max license.
Missing a frontier release? Add a model (editors)