DeepSeek-V4.1-Flash
DeepSeek-V4.1-Flash is DeepSeek's September 10, 2026 multimodal Mixture-of-Experts model and the smallest member of its new architecture family.
Provider
DeepSeek
Model family
DeepSeek
Multimodal MoE LLM
Cost tier
Flash
Status
Current
Release Sep 10, 2026
Why teams choose it
Call deepseek-flash. Legacy names deepseek-v4-flash and deepseek-v4-flash-vision-exp sti…
ll route here and bill at Flash rates.
Official peak/off-peak list (per 1M tokens): cache-hit $0.003/$0.006, cache-miss $0.15/$0.30, output $0.60/$1.20
Off-peak is half of peak. Peak is 01:00–04:00 and 06:00–10:00 UTC, Monday–Friday.
From 04:00 UTC on September 14, 2026, deepseek-v4-pro requests route to V4.1-Flash at Fl…
ash rates until V4.1-Pro launches.
Tradeoffs to know
- Peak/off-peak hours are UTC weekday windows; weekend traffic is off-peak.
- Vendor benches use DeepSeek Harness and max reasoning_effort=100; rematch on your scaffold.
When not to use this
- V4.1-Pro is announced, not released. Do not invent a Pro SKU yet.
- Self-hosting outcomes depend on hardware, quantization, and ops maturity—budget time beyond swapping an API hostname.
- May demand more instrumentation than SaaS-managed APIs to duplicate latency, failover, and support guarantees.
Technical specs
- Inputs
- text, image
- Outputs
- text
- Capabilities
- reasoning, agentic coding, tool calling, native vision, long context, controllable thinking
- License
- MIT
- Model string
deepseek-v4-1-flash
Benchmarks
{
"api_id": "deepseek-flash",
"source": "https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash",
"active_params": "8B prefill / 16B decode",
"context_tokens": 1000000,
"backbone_params": "552B",
"vendor_reported": true
}DeepSeek family lineup
Current models
Compare with
Available through
Verified subscription plans that list this model.
DeepSeek-V4.1-Flash FAQ
What is DeepSeek-V4.1-Flash?
DeepSeek-V4.1-Flash is DeepSeek's September 10, 2026 multimodal Mixture-of-Experts model and the smallest member of its new architecture family. The live API ID is deepseek-flash. Official materials describe a 552B-parameter Causal Encoder–Decoder MoE that activates 8B parameters on input and 16B on output, native ima...
When does DeepSeek-V4.1-Flash fit best?
High-volume agentic coding at Flash cost
What should teams watch out for with DeepSeek-V4.1-Flash?
V4.1-Pro is announced, not released. Do not invent a Pro SKU yet.
Explore next
Models, tools, and comparisons that connect to this reference.