Structured cards
Model database
Filter by provider, architecture family, or full-text search across descriptions.
Frontier models
Verified flagships and recent launches across major providers—scroll sideways for the full shelf.
All models
Filter and paginate the full catalog. Tabs control lifecycle scope.
Gemini 1.0 Pro
LegacyGemini 1.0 Pro represents Google’s first broadly marketed Gemini-era general model for text and basic multimodal tasks on Vertex and consumer surfaces. New projects should prefer 1.5+ generations unless constrained by legacy integrations—verify availability.
Cohere
Command R
LegacyCommand R is Cohere’s earlier RAG-oriented model line preceding Command R+, focused on grounded generation with connectors and multilingual enterprise search. Useful when comparing tiered Cohere stacks or maintaining legacy integrations.
OpenAI
GPT-4 Turbo
LegacyGPT-4 Turbo is a widely deployed GPT-4-class chat model with a large context window on the OpenAI API, aimed at long-document workflows, retrieval bundles, and production assistants that do not require GPT-4o’s multimodal stack. It remains a common baseline for cost/quality tradeoffs.
Anthropic
Claude 3 Haiku
LegacyClaude 3 Haiku is the fast Claude 3-era model for simple tasks and high throughput. Prefer 3.5 Haiku when available for better quality at similar latency targets—confirm SKUs on your cloud path.
OpenAI
GPT-4o mini
LegacyGPT-4o mini is a cost-optimized GPT-4o-family model for high-volume chat, moderation, and routing layers where frontier quality is unnecessary. It supports multimodal inputs on supported API surfaces and is often used as a fast first pass before escalating to larger models.
Gemini 2.0 Flash
LegacyGemini 2.0 Flash is Google’s efficiency-oriented multimodal model generation aimed at fast agentic and interactive experiences. Capabilities and naming evolve—validate against the current Gemini API reference for tool use and context limits.
01.AI
Yi-34B
LegacyCatalog entry for this named release; see the provider’s official documentation for modalities, pricing, and context limits.
Alibaba
Qwen 2
LegacyCatalog entry for this named release; see the provider’s official documentation for modalities, pricing, and context limits.
Microsoft
Phi-3 Mini
LegacyCatalog entry for this named release; see the provider’s official documentation for modalities, pricing, and context limits.
Meta
Llama 3 8B
LegacyCatalog entry for this named release; see the provider’s official documentation for modalities, pricing, and context limits.
Meta
Llama 3 70B
LegacyCatalog entry for this named release; see the provider’s official documentation for modalities, pricing, and context limits.
OpenAI
GPT-4.1 nano
LegacyCatalog entry for this named release; see the provider’s official documentation for modalities, pricing, and context limits.
TII
Falcon 180B
LegacyCatalog entry for this named release; see the provider’s official documentation for modalities, pricing, and context limits.
Anthropic
Claude Opus 4.5
LegacyCatalog entry for this named release; see the provider’s official documentation for modalities, pricing, and context limits.
Anthropic
Claude Opus 4.6
LegacyCatalog entry for this named release; see the provider’s official documentation for modalities, pricing, and context limits.
Anthropic
Claude Sonnet 4.5
LegacyCatalog entry for this named release; see the provider’s official documentation for modalities, pricing, and context limits.
OpenAI
DALL·E 3
DeprecatedDALL·E 3 is OpenAI’s instruction-aligned image generation model exposed via the Images API, emphasizing prompt adherence and safety classifiers for consumer and enterprise creative workflows. It targets marketing visuals, product mockups, and storyboarding rather than photorealistic deception.
OpenAI
o1-mini
DeprecatedA smaller, faster o1-class model for STEM-style reasoning where full o1 latency or cost is prohibitive. Use it when you need better chain-of-thought style behavior than GPT-4o mini but not full o1 depth.
Recommended
Top current models by information quality score — good defaults when you are not sure where to start.
Sarvam AI
Sarvam 30B
CurrentLatestSarvam 30B is a 30B parameter Mixture-of-Experts chat and reasoning model from Sarvam AI, optimized for Indian languages, real-time conversation, high-throughput voice-agent pipelines, coding, and practical deployment. Sarvam documents 2.4B active parameters per token, 16T tokens of pre-training data, a 64K context window, Grouped Query Attention, Apache 2.0 open weights, and OpenAI-compatible chat completions.
MiniMax
MiniMax M3
CurrentLatestMiniMax M3 is a June 2026 open-weight multimodal model for coding, agentic workflows, computer use, and long-context work. MiniMax documents a one-million-token context window, native image and video understanding, and deployment through hosted or downloadable model paths.
Alibaba Qwen
Qwen3.8-2.4T-A95B
CurrentLatestQwen3.8-2.4T-A95B is Qwen's downloadable 2.4-trillion-parameter sparse mixture-of-experts model with 95B activated parameters. The official model card documents text-only input and output, mandatory thinking, 262,144 native context extensible to roughly 1.01M, configurable reasoning effort, and a dedicated Qwen3.8-Max license.
Missing a frontier release? Add a model (editors)