Inco AI launches Day-0 support for GLM 5.3
GLM 5.3 — Cumulative Serving Speedup
End-to-end serving throughput
Overview
GLM 5.3 arrives today with day-0 serving support from Inco AI on TokenRouter compute. The release pairs Z.ai's new model with Inco AI's core inference optimization technology:
- A DFlash 2 checkpoint for speculative decoding.
- An NVFP4 checkpoint for efficient Blackwell inference with native accuracy.
- Inco Engine for end-to-end inference performance.
As an official day-0 partner of Z.ai, Inco AI is releasing both checkpoints alongside the model endpoint. Together, the checkpoints and Inco Engine deliver up to 4.42× throughput versus the native FP8 checkpoint with autoregressive decoding at concurrency 1.
DFlash 2 for GLM 5.3
DFlash 2 is a parallel drafter for speculative decoding: it predicts candidate tokens in one pass and lets the target model verify them as a block.
Acceptance length (AL) measures how many tokens each draft–verify cycle yields, including the verifier's next token. A longer accepted block means fewer full target-model passes for the same output. DFlash 2 improves AL by keeping the parallel draft design while selecting a more coherent path through each position's candidates.
Acceptance length is only half of the serving result: a drafter also adds work to each cycle. For GLM 5.3, we therefore report end-to-end decode throughput separately, comparing the model's native MTP path and DFlash 2 against autoregressive decoding. At concurrency 1, DFlash 2 reaches 383.3 output tok/s on MATH-500, 366.6 on GSM8K, and 363.7 on HumanEval.
GLM 5.3 — DFlash 2 performance
Acceptance length
Higher is betterDecode speedup
Concurrency 1NVFP4 for Blackwell
We are also releasing an NVFP4 checkpoint for efficient Blackwell deployment, with accuracy matching the native FP8 checkpoint across the evaluation suite.
| Precision | GPQA Diamond | AIME 2025 | MATH-500 | HLE | AA-LCR |
|---|---|---|---|---|---|
| FP8 | 91.1 | 94.3 | 95.6 | 35.9 | 73.6 |
| NVFP4 | 91.2 | 95.1 | 95.2 | 35.2 | 73.0 |
Get access
Try GLM 5.3 in the browser or connect through the Inco API.
Included in preview
- Browser playground
- OpenAI and Anthropic API dialects
- Streaming responses
- 1M-token context window
List pricing
per 1M tokens- Input
- $1.40
- Cached input
- $0.26
- Output
- $4.40
Get checkpoints
Compute for GLM 5.3
Compute for this project came from TokenRouter. The Blackwell capacity used to train the DFlash 2 drafter and produce the NVFP4 checkpoint ran on TokenRouter clusters. The day-0 GLM 5.3 endpoint is served from the same pool and reachable with an existing TokenRouter key.
Get updates
One email when we ship something new.
We will never share your email address.
