GLM-5.3-Flash
GLM-5.3-Flash is Z.ai's first natively multimodal GLM-5 model: a 320B-parameter mixture-of-experts checkpoint with 18B active parameters, hybrid sparse and linear attention, and MIT-licensed weights.
Provider
Z.ai
Model family
Z.ai GLM
Open-weight multimodal MoE LLM
Cost tier
Flash
Status
Current
Release Aug 26, 2026
Why teams choose it
Z.ai says the model was previously previewed anonymously as ox-alpha on OpenCode and OpenRouter.
Z.ai says the model was previously previewed anonymously as ox-alpha on OpenCode and OpenRouter.
Official architecture notes 320B total / 18B active parameters
hybrid sparse plus linear attention, and about 1M-token long context.
This Flash SKU is distinct from hosted GLM-5.3: it is natively multimodal and ships MIT weights.
This Flash SKU is distinct from hosted GLM-5.3: it is natively multimodal and ships MIT weights.
Tradeoffs to know
- Published coding, agent, and vision scores are vendor-reported and harness-sensitive.
- Self-hosting a 320B MoE still requires substantial GPU memory, a supported runtime, and quantization choices.
- List pricing, launch discounts, and Coding Plan quotas change; confirm current Z.ai terms before production use.
When not to use this
- Self-hosting outcomes depend on hardware, quantization, and ops maturity—budget time beyond swapping an API hostname.
- May demand more instrumentation than SaaS-managed APIs to duplicate latency, failover, and support guarantees.
- Benchmark prompts and regressions continuously before rewriting entire routing tables around weights.
Technical specs
- Inputs
- text, image, video
- Outputs
- text
- Capabilities
- coding, agentic workflows, multimodal reasoning, visual coding, computer use, long-context, open-weight serving
- License
- MIT
- Model string
glm-5-3-flash
Benchmarks
{
"source": "https://z.ai/blog/glm-5.3-flash",
"deep_swe_v1_1": 63.4,
"vendor_reported": true,
"terminal_bench_2_1": 84.3,
"toolathlon_verified": 78.4,
"aa_intelligence_index": 57,
"automationbench_v1_0_6": 48.8
}Z.ai GLM family lineup
Current models
Previous versions
Compare with
GLM-5.3-Flash FAQ
What is GLM-5.3-Flash?
GLM-5.3-Flash is Z.ai's first natively multimodal GLM-5 model: a 320B-parameter mixture-of-experts checkpoint with 18B active parameters, hybrid sparse and linear attention, and MIT-licensed weights. Z.ai documents API access as glm-5.3-flash, Coding Plan and ZCode availability, and Hugging Face weights for SGLang, vL...
When does GLM-5.3-Flash fit best?
Cost-sensitive coding agents
What should teams watch out for with GLM-5.3-Flash?
Published coding, agent, and vision scores are vendor-reported and harness-sensitive.
Explore next
Models, tools, and comparisons that connect to this reference.