GenAIWiki

Qwen3.8-2.4T-A95B

CurrentLatest

Qwen3.8-2.4T-A95B is Qwen's downloadable 2.4-trillion-parameter sparse mixture-of-experts model with 95B activated parameters.

Provider

Alibaba Qwen

Model family

Alibaba Qwen

Open-weight sparse MoE reasoning LLM

Cost tier

Open Flagship

Status

Current

Release Aug 12, 2026

Why teams choose it

🧠

The checkpoint activates about 95B of 2.4T total parameters per token and is a distinct…

product from hosted Qwen3.8-Max.

📎

The released checkpoint is text-only and always uses thinking; vision and non-thinking s…

upport belong to the hosted Max service.

⚙️

Native context is 262,144 tokens, with documented extension to roughly 1.01M under suppo…

rted serving configurations.

Tradeoffs to know

  • The 2.4T checkpoint has exceptional storage, memory, networking, and serving requirements despite sparse activation.
  • Its custom Qwen3.8-Max License is not Apache 2.0; review attribution and commercial-service obligations.
  • Model-card capabilities and evaluations are vendor-reported and sensitive to serving stack, precision, and reasoning configuration.

When not to use this

  • Self-hosting outcomes depend on hardware, quantization, and ops maturity—budget time beyond swapping an API hostname.
  • May demand more instrumentation than SaaS-managed APIs to duplicate latency, failover, and support guarantees.
  • Benchmark prompts and regressions continuously before rewriting entire routing tables around weights.

Technical specs

Inputs
text
Outputs
text
Capabilities
reasoning, coding, research, long-horizon agents, long-context, configurable reasoning, local deployment
License
Qwen3.8-Max License
Model string
qwen-3-8-2-4t-a95b

Benchmarks

{
  "source": "https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B",
  "native_context": 262144,
  "vendor_reported": true,
  "total_parameters": "2.4T",
  "activated_parameters": "95B"
}

Alibaba Qwen family lineup


Compare with

Qwen3.8-2.4T-A95B FAQ

What is Qwen3.8-2.4T-A95B?

Qwen3.8-2.4T-A95B is Qwen's downloadable 2.4-trillion-parameter sparse mixture-of-experts model with 95B activated parameters. The official model card documents text-only input and output, mandatory thinking, 262,144 native context extensible to roughly 1.01M, configurable reasoning effort, and a dedicated Qwen3.8-Max...

When does Qwen3.8-2.4T-A95B fit best?

Large-scale self-hosted reasoning

What should teams watch out for with Qwen3.8-2.4T-A95B?

The 2.4T checkpoint has exceptional storage, memory, networking, and serving requirements despite sparse activation.

Explore next

Models, tools, and comparisons that connect to this reference.