GenAIWiki

DeepSeek-V4.1-Flash

CurrentLatest

DeepSeek-V4.1-Flash is DeepSeek's September 10, 2026 multimodal Mixture-of-Experts model and the smallest member of its new architecture family.

Provider

DeepSeek

Model family

DeepSeek

Multimodal MoE LLM

Cost tier

Flash

Status

Current

Release Sep 10, 2026

Why teams choose it

🧠

Call deepseek-flash. Legacy names deepseek-v4-flash and deepseek-v4-flash-vision-exp sti…

ll route here and bill at Flash rates.

📎

Official peak/off-peak list (per 1M tokens): cache-hit $0.003/$0.006, cache-miss $0.15/$0.30, output $0.60/$1.20

Off-peak is half of peak. Peak is 01:00–04:00 and 06:00–10:00 UTC, Monday–Friday.

⚙️

From 04:00 UTC on September 14, 2026, deepseek-v4-pro requests route to V4.1-Flash at Fl…

ash rates until V4.1-Pro launches.

Tradeoffs to know

  • Peak/off-peak hours are UTC weekday windows; weekend traffic is off-peak.
  • Vendor benches use DeepSeek Harness and max reasoning_effort=100; rematch on your scaffold.

When not to use this

  • V4.1-Pro is announced, not released. Do not invent a Pro SKU yet.
  • Self-hosting outcomes depend on hardware, quantization, and ops maturity—budget time beyond swapping an API hostname.
  • May demand more instrumentation than SaaS-managed APIs to duplicate latency, failover, and support guarantees.

Technical specs

Inputs
text, image
Outputs
text
Capabilities
reasoning, agentic coding, tool calling, native vision, long context, controllable thinking
License
MIT
Model string
deepseek-v4-1-flash

Benchmarks

{
  "api_id": "deepseek-flash",
  "source": "https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash",
  "active_params": "8B prefill / 16B decode",
  "context_tokens": 1000000,
  "backbone_params": "552B",
  "vendor_reported": true
}

DeepSeek family lineup


Compare with

Available through

Verified subscription plans that list this model.

DeepSeek-V4.1-Flash FAQ

What is DeepSeek-V4.1-Flash?

DeepSeek-V4.1-Flash is DeepSeek's September 10, 2026 multimodal Mixture-of-Experts model and the smallest member of its new architecture family. The live API ID is deepseek-flash. Official materials describe a 552B-parameter Causal Encoder–Decoder MoE that activates 8B parameters on input and 16B on output, native ima...

When does DeepSeek-V4.1-Flash fit best?

High-volume agentic coding at Flash cost

What should teams watch out for with DeepSeek-V4.1-Flash?

V4.1-Pro is announced, not released. Do not invent a Pro SKU yet.

Explore next

Models, tools, and comparisons that connect to this reference.