Structured cards
Model database
Filter by provider, architecture family, or full-text search across descriptions.
Frontier models
Verified flagships and recent launches across major providers—scroll sideways for the full shelf.
All models
Filter and paginate the full catalog. Tabs control lifecycle scope.
Gemini 1.5 Pro
LegacyGoogle DeepMind Gemini 1.5 Pro targets long-context multimodal workloads—large effective context for retrieval-heavy document pipelines, plus image, audio, and video inputs on supported surfaces. It is often paired with Vertex AI or the Gemini API for enterprise workloads on GCP.
DeepSeek
DeepSeek-V3
LegacyDeepSeek-V3 is a large-scale language model family noted for strong coding and math performance under open or research-friendly terms (verify the exact license for your deployment). Teams adopt it for cost-sensitive research, self-hosted inference, or comparison against frontier APIs.
Anthropic
Claude 3.5 Sonnet
LegacyAnthropic’s balanced Sonnet-tier model tuned for long-context reasoning, careful instruction following, and strong performance on coding and analysis workloads. It is a common enterprise choice on the Anthropic API and on AWS Bedrock when teams need large context for RAG and document review.
Mistral AI
Mistral Large 2
LegacyMistral’s frontier-class multilingual model emphasizing JSON adherence, agent-friendly behavior, and competitive reasoning within the Mistral API ecosystem. European teams often evaluate it for GDPR-adjacent deployment patterns alongside US-hosted alternatives.
OpenAI
GPT-5.4
LegacyOpenAI's GPT-5.4 model, documented in the official OpenAI API model guide as part of the current GPT-5 family below the GPT-5.5 flagship lane.
OpenAI
GPT-4o
LegacyGPT-4o is an OpenAI multimodal model that accepts text and image inputs and produces text. It supports streaming, function calling, Structured Outputs, fine-tuning, and predicted outputs for vision-heavy assistants and structured extraction workflows.
xAI
Grok-2
LegacyGrok-2 is xAI’s flagship chat model positioned for real-time knowledge integrations and high-throughput conversational products on xAI’s API. Availability and pricing evolve—treat capabilities as vendor-specific.
Alibaba
Qwen 2.5 72B Instruct
LegacyQwen 2.5 72B Instruct is a large multilingual open-weights model from Alibaba’s Qwen family with strong coding and general chat performance. Common in APAC deployments and on Hugging Face inference endpoints—check license terms for commercial use.
Anthropic
Claude Sonnet 4.6
LegacyAnthropic's Sonnet-tier model documented as the best combination of speed and intelligence in the Claude model overview.
OpenAI
o1
LegacyOpenAI’s o1 series emphasizes extended internal reasoning before answering—useful for competition-style math, complex debugging, and multi-step planning where latency is acceptable. It behaves differently from standard chat models: tune prompts for chain-of-thought style tasks and measure time-to-first-token.
Gemini 2.5 Flash-Lite
LegacyGoogle's fastest and most budget-friendly multimodal model in the Gemini 2.5 family, according to the Gemini API model documentation.
DeepSeek
DeepSeek-V3.2
LegacyDeepSeek's documented successor to the V3.2 experimental line, positioned in official DeepSeek API news as live on app, web, and API.
OpenAI
GPT-5.4 mini
LegacyOpenAI's smaller GPT-5.4 mini model, documented in the official OpenAI API model guide for lower-latency or lower-cost GPT-5 family workloads.
OpenAI
GPT-5.4 nano
LegacyOpenAI's smallest GPT-5.4 nano model, documented in the official OpenAI API model guide for very low-latency or economical GPT-5 family routing.
Gemini 1.5 Flash
LegacyGemini 1.5 Flash targets low-latency, cost-efficient multimodal chat and retrieval workloads on the Gemini API and Vertex AI. It keeps much of the long-context family behavior with faster responses for interactive apps.
Anthropic
Claude 3.5 Haiku
LegacyClaude 3.5 Haiku is Anthropic’s fast, cost-efficient tier for high-volume classification, routing, and simple chat. It targets latency-sensitive paths and agent pre-processing before escalating to Sonnet-class models.
Microsoft
Phi-3 Medium
LegacyPhi-3 Medium is a compact instruct model aimed at strong quality per parameter for on-device and cost-sensitive cloud inference. It competes with other SLMs on coding and reasoning benchmarks—validate on your domain prompts.
Meta
Llama 3.2 1B Instruct
LegacyLlama 3.2 1B Instruct is among the smallest Llama instruct checkpoints for extreme latency and footprint constraints. Use for routing, tagging, and toy assistants—not for complex reasoning without retrieval augmentation.
Mistral AI
Mixtral 8x7B Instruct
LegacyMixtral 8x7B Instruct is a sparse mixture-of-experts open model noted for strong quality per active parameter and efficient inference vs dense models of similar capability. Widely hosted on inference clouds and self-hosted stacks.
OpenAI
GPT-3.5 Turbo
LegacyGPT-3.5 Turbo is a long-standing cost-efficient chat model family on the OpenAI API for simple assistants, classification, and legacy integrations. Many teams still use it for non-critical paths or as a fallback when newer models are rate-limited.
Anthropic
Claude 3 Opus
LegacyClaude 3 Opus was Anthropic’s highest-capability Claude 3-era model for difficult reasoning, nuanced writing, and complex analysis before later Sonnet generations. Teams still reference it for historical benchmarks and legacy deployments—verify current availability in API and Bedrock model lists.
Anthropic
Claude 3 Sonnet
LegacyClaude 3 Sonnet balanced cost and capability in the Claude 3 generation—useful for general assistants and document workflows where Opus was unnecessary. New deployments should compare against Claude 3.5 Sonnet for pricing and quality.
Meta
Llama 3.2 3B Instruct
LegacyLlama 3.2 3B Instruct is a compact instruct model in Meta’s 3.2 generation aimed at mobile and edge scenarios with multilingual support on supported checkpoints. Verify hardware targets and license terms for your distribution channel.
xAI
Grok-3
LegacyGrok-3 represents xAI’s newer generation aimed at stronger reasoning and tool use versus Grok-2. Capabilities and rollout are version-specific—validate against xAI documentation for your account tier.
Recommended
Top current models by information quality score — good defaults when you are not sure where to start.
Sarvam AI
Sarvam 30B
CurrentLatestSarvam 30B is a 30B parameter Mixture-of-Experts chat and reasoning model from Sarvam AI, optimized for Indian languages, real-time conversation, high-throughput voice-agent pipelines, coding, and practical deployment. Sarvam documents 2.4B active parameters per token, 16T tokens of pre-training data, a 64K context window, Grouped Query Attention, Apache 2.0 open weights, and OpenAI-compatible chat completions.
MiniMax
MiniMax M3
CurrentLatestMiniMax M3 is a June 2026 open-weight multimodal model for coding, agentic workflows, computer use, and long-context work. MiniMax documents a one-million-token context window, native image and video understanding, and deployment through hosted or downloadable model paths.
Alibaba Qwen
Qwen3.8-2.4T-A95B
CurrentLatestQwen3.8-2.4T-A95B is Qwen's downloadable 2.4-trillion-parameter sparse mixture-of-experts model with 95B activated parameters. The official model card documents text-only input and output, mandatory thinking, 262,144 native context extensible to roughly 1.01M, configurable reasoning effort, and a dedicated Qwen3.8-Max license.
Missing a frontier release? Add a model (editors)