Z.AI · model

GLM 5.3 Flash

$0.075/1M in · $0.25/1M out · 1.3M context. That's cheaper than 89% of the 389 models we track, by output price. Here's what it costs, how its price has moved, and where it fits.

output price · per 1M tokens
$0.25
input $0.075/1M · routed
Context
1.3M
Input / 1M
$0.075
Output / 1M
$0.25
Modality
text + image + video
Provider
Z.AI
Tokenizer
Other

Snapshot · source: OpenRouter ↗ · how we measure →

cost at a typical workload
$47.50 / mo

For 1,500 input + 500 output tokens across 200,000 requests/month. That's 5.5× the cheapest tracked option (Mistral Nemo, $8.70/mo).

Tune the workload in the calculator →

output price history
output / 1M

Price history is accruing — the archive began 2026-08-02 and captures a snapshot every other day. Track it on the history page →

Where it fits

  • High-volume, cost-sensitive work — at $0.25/1M output it sits in the cheapest quarter of tracked models.
  • Long inputs — a 1.3M-token window holds roughly 1,966 pages of text at once.
  • Multimodal prompts — accepts image input alongside text.
  • Input-heavy jobs like summarization and RAG — input is cheap at $0.075/1M, and these read far more than they write.

Watch for

  • Priced from OpenRouter's routed market, not a first-party list price — availability and rate can shift without notice.

These notes are derived from price, context and modality — structural facts, not measured quality. Measured cost-to-finish-a-task lives on the real-cost index.

Related models

See how the leading models compare head-to-head →

How to read GLM 5.3 Flash's pricing

Two numbers decide most of the bill: $0.075/1M for input (everything you send — prompt, context, attachments) and $0.25/1M for output (everything it generates, including hidden reasoning tokens). Output is priced 3.3× the input rate here, so the shape of your workload — how much it reads versus writes — matters as much as the headline figure. This row is OpenRouter's routed market price; it can move as providers and routing change.

Frequently asked questions

How much does GLM 5.3 Flash cost per million tokens?

GLM 5.3 Flash is priced at $0.075 per 1M input tokens and $0.25 per 1M output tokens, from OpenRouter's routed market price as of the 2026-08-28 snapshot. Output is the figure that usually drives the bill. Enter your own token volumes in the cost calculator for a monthly estimate.

What is GLM 5.3 Flash's context window?

GLM 5.3 Flash has a 1.3M-token context window — roughly 1,966 pages of text. Prompt, attachments, conversation and the model's own output all share that budget, and every token you send is billed at the input rate.

Related