GenAIWiki

GLM-5.3-Flash

CurrentLatest

GLM-5.3-Flash is Z.ai's first natively multimodal GLM-5 model: a 320B-parameter mixture-of-experts checkpoint with 18B active parameters, hybrid sparse and linear attention, and MIT-licensed weights.

Provider

Z.ai

Model family

Z.ai GLM

Open-weight multimodal MoE LLM

Cost tier

Flash

Status

Current

Release Aug 26, 2026

Why teams choose it

🧠

Z.ai says the model was previously previewed anonymously as ox-alpha on OpenCode and OpenRouter.

Z.ai says the model was previously previewed anonymously as ox-alpha on OpenCode and OpenRouter.

📎

Official architecture notes 320B total / 18B active parameters

hybrid sparse plus linear attention, and about 1M-token long context.

⚙️

This Flash SKU is distinct from hosted GLM-5.3: it is natively multimodal and ships MIT weights.

This Flash SKU is distinct from hosted GLM-5.3: it is natively multimodal and ships MIT weights.

Tradeoffs to know

  • Published coding, agent, and vision scores are vendor-reported and harness-sensitive.
  • Self-hosting a 320B MoE still requires substantial GPU memory, a supported runtime, and quantization choices.
  • List pricing, launch discounts, and Coding Plan quotas change; confirm current Z.ai terms before production use.

When not to use this

  • Self-hosting outcomes depend on hardware, quantization, and ops maturity—budget time beyond swapping an API hostname.
  • May demand more instrumentation than SaaS-managed APIs to duplicate latency, failover, and support guarantees.
  • Benchmark prompts and regressions continuously before rewriting entire routing tables around weights.

Technical specs

Inputs
text, image, video
Outputs
text
Capabilities
coding, agentic workflows, multimodal reasoning, visual coding, computer use, long-context, open-weight serving
License
MIT
Model string
glm-5-3-flash

Benchmarks

{
  "source": "https://z.ai/blog/glm-5.3-flash",
  "deep_swe_v1_1": 63.4,
  "vendor_reported": true,
  "terminal_bench_2_1": 84.3,
  "toolathlon_verified": 78.4,
  "aa_intelligence_index": 57,
  "automationbench_v1_0_6": 48.8
}

Z.ai GLM family lineup


Compare with

GLM-5.3-Flash FAQ

What is GLM-5.3-Flash?

GLM-5.3-Flash is Z.ai's first natively multimodal GLM-5 model: a 320B-parameter mixture-of-experts checkpoint with 18B active parameters, hybrid sparse and linear attention, and MIT-licensed weights. Z.ai documents API access as glm-5.3-flash, Coding Plan and ZCode availability, and Hugging Face weights for SGLang, vL...

When does GLM-5.3-Flash fit best?

Cost-sensitive coding agents

What should teams watch out for with GLM-5.3-Flash?

Published coding, agent, and vision scores are vendor-reported and harness-sensitive.

Explore next

Models, tools, and comparisons that connect to this reference.