Structured cards
Model database
Filter by provider, architecture family, or full-text search across descriptions.
Frontier models
Verified flagships and recent launches across major providers—scroll sideways for the full shelf.
All models
Filter and paginate the full catalog. Tabs control lifecycle scope.
DeepSeek
DeepSeek-V4-Pro
CurrentLatestDeepSeek-V4-Pro is DeepSeek's V4 model for high-capability reasoning and agentic coding, available through DeepSeek's OpenAI-compatible and Anthropic-compatible API surfaces.
OpenAI
GPT-5.6 Luna
CurrentLatestOpenAI's GPT-5.6 Luna is the cost-sensitive GPT-5.6 tier for high-volume workloads that still need current GPT-5.6 behavior, vision input, structured outputs, and tool support.
Meta
Llama 3.1 405B Instruct
CurrentLatestMeta’s largest open-weights instruct checkpoint in the Llama 3.1 family, aimed at strong reasoning and coding quality with a permissive license for research and customization. It is typically served on dedicated GPU clusters or via partners (cloud inference, on-prem) rather than a single vendor API.
DeepSeek
DeepSeek-R1
CurrentLatestDeepSeek-R1 is a reasoning-focused model family emphasizing chain-of-thought style behavior for math, code, and structured problem solving. Deployment options include API and open-weight variants—verify licensing and hosting constraints for your region.
Anysphere
Cursor Composer 2.5
CurrentLatestCursor Composer 2.5 is Anysphere's price-efficient first-party coding model for long-running agentic tasks inside Cursor, with improved sustained work and instruction following over Composer 2.
Cohere
Command R+
CurrentLatestCohere’s enterprise-oriented Command R+ emphasizes retrieval-grounded answers and tool orchestration patterns for business data. It targets teams building RAG-heavy assistants where citation-style behavior and connector patterns matter more than raw chat novelty.
Gemini 3.1 Flash-Lite
CurrentLatestGemini 3.1 Flash-Lite is Google's stable Gemini 3-series workhorse model for cost-efficient, high-volume multimodal workloads.
Anthropic
Claude Haiku 4.5
CurrentLatestClaude Haiku 4.5 is Anthropic's fastest current Claude model with near-frontier intelligence for high-volume and latency-sensitive workloads.
DeepSeek
DeepSeek-V4-Flash
CurrentLatestDeepSeek-V4-Flash is DeepSeek's faster and more economical V4 model, supporting thinking and non-thinking modes through the current DeepSeek API.
Microsoft AI
MAI-Code-1-Flash
CurrentLatestMicrosoft AI's agentic coding model in the MAI family, announced for fast code editing, debugging, and tool-driven developer workflows.
Mistral AI
Mistral Large 3
CurrentLatestMistral's open-weight general-purpose multimodal model listed in official Mistral model documentation.
Microsoft AI
MAI-Image-2.5
CurrentLatestMicrosoft AI's MAI image model for generation, editing, and visual content workflows, announced as part of the June 2026 MAI model release.
Microsoft AI
MAI-Voice-2
CurrentLatestMicrosoft AI's voice generation model in the MAI family, announced for natural text-to-speech and voice experiences.
Microsoft AI
MAI-Transcribe-1.5
CurrentLatestMicrosoft AI's speech-to-text model in the MAI family, announced for fast, accurate transcription across product surfaces.
Microsoft AI
MAI-Image-2.5-Flash
CurrentLatestMicrosoft AI's faster MAI image variant, announced for lower-latency image generation and editing workflows.
Mistral AI
Mistral Small 3
CurrentLatestMistral Small 3 is Mistral’s efficiency tier for fast, affordable chat and tool use at high QPS—positioned between tiny open models and Mistral Large. Exact naming and versioning appear in Mistral’s API catalog; pin versions in production.
Meta
Llama 3.1 70B Instruct
CurrentLatestLlama 3.1 70B Instruct is a mid-size open-weights instruct model balancing quality and deployability on a single large GPU or small multi-GPU nodes. Common for private assistants, on-prem pilots, and fine-tunes where 405B is impractical.
OpenAI
Whisper large-v3
CurrentLatestWhisper large-v3 is OpenAI’s ASR model for transcription and translation across many languages, with strong robustness to accents and noise. It is commonly self-hosted or used via API partners; latency depends heavily on hardware and chunking strategy.
Snowflake
Snowflake Arctic
CurrentLatestSnowflake Arctic is an enterprise-oriented open model emphasizing efficient training recipes and SQL-adjacent enterprise tasks inside the Snowflake ecosystem. It targets teams that want LLM features colocated with governed data in Snowflake Cortex.
NVIDIA
NVIDIA Nemotron-4 340B
CurrentLatestNVIDIA Nemotron-4 340B is a large open-weights model suite aimed at enterprise and research users who train and serve on NVIDIA stacks (NeMo, NGC). It targets GPU-native teams that need customizable checkpoints with NVIDIA-optimized tooling.
Mistral AI
Mistral 7B Instruct v0.3
CurrentLatestMistral 7B Instruct is a compact dense model that popularized efficient open-weight chat quality at small scale. It remains a baseline for fine-tunes and on-prem pilots where 13B+ models are too heavy.
OpenAI
o3-mini
CurrentLatestCompact reasoning-focused model in OpenAI’s o-series line aimed at strong STEM and coding performance with lower cost than full o3. Intended for developers who want reasoning without always paying flagship prices—confirm exact API availability and snapshot names in OpenAI docs.
OpenAI
text-embedding-3-large
CurrentLatesttext-embedding-3-large produces high-dimensional text embeddings for semantic search, clustering, and classification. Teams pair it with pgvector or SaaS vector DBs for RAG; output dimensions can be reduced with tradeoffs described in OpenAI documentation.
Microsoft
Phi-4
CurrentLatestPhi-4 is Microsoft Research’s small language model line focused on strong reasoning per parameter for on-device and low-cost cloud scenarios. Deployment often happens via Azure AI or Hugging Face hubs—confirm license for your channel.
Recommended
Top current models by information quality score — good defaults when you are not sure where to start.
Sarvam AI
Sarvam 30B
CurrentLatestSarvam 30B is a 30B parameter Mixture-of-Experts chat and reasoning model from Sarvam AI, optimized for Indian languages, real-time conversation, high-throughput voice-agent pipelines, coding, and practical deployment. Sarvam documents 2.4B active parameters per token, 16T tokens of pre-training data, a 64K context window, Grouped Query Attention, Apache 2.0 open weights, and OpenAI-compatible chat completions.
MiniMax
MiniMax M3
CurrentLatestMiniMax M3 is a June 2026 open-weight multimodal model for coding, agentic workflows, computer use, and long-context work. MiniMax documents a one-million-token context window, native image and video understanding, and deployment through hosted or downloadable model paths.
Alibaba Qwen
Qwen3.8-2.4T-A95B
CurrentLatestQwen3.8-2.4T-A95B is Qwen's downloadable 2.4-trillion-parameter sparse mixture-of-experts model with 95B activated parameters. The official model card documents text-only input and output, mandatory thinking, 262,144 native context extensible to roughly 1.01M, configurable reasoning effort, and a dedicated Qwen3.8-Max license.
Missing a frontier release? Add a model (editors)