Z.ai: GLM 5.3 FlashX

Video input Image input Text input Text output
Author's Description

GLM-5.3-FlashX is the high-speed variant of Z.ai's GLM-5.3-Flash, a native multimodal model delivering inference speeds of up to 200 tokens/s. Built on the same hybrid sparse and linear attention architecture...

Key Specifications
Cost
$$$$
Context
1M
Released
Sep 18, 2026
Speed
Ability
Reliability
Supported Parameters

This model supports the following parameters:

Top P Max Tokens Temperature Tool Choice Response Format Include Reasoning Tools Reasoning
Features

This model supports the following features:

Tools Response Format Reasoning
Performance Summary

Z.ai's GLM 5.3 FlashX, a high-speed multimodal model, demonstrates competitive performance across various benchmarks. Its speed ranking places it in the 49th percentile, indicating generally competitive response times, while its pricing is also competitive, ranking in the 43rd percentile. A standout feature is its exceptional reliability, achieving a 100% success rate across all 8 benchmarks, signifying consistent and dependable operation with minimal technical failures. The model exhibits strong performance in several key areas. It achieved perfect accuracy in both General Knowledge and Reasoning, with these benchmarks also highlighting it as the most accurate model at its price point and among models of similar speed. Its Coding and Mathematics capabilities are also impressive, scoring 97.0% accuracy and ranking in the 97th and 96th percentiles respectively. Instruction Following and Email Classification also show high accuracy at 78.0% and 99.0%. The model demonstrates a high degree of non-hallucination with 98.0% accuracy. While its Ethics score of 99.0% is solid, it falls in the 51st percentile, indicating room for improvement relative to other models in this specific area. Overall, GLM 5.3 FlashX is a highly reliable and accurate model, particularly strong in knowledge, reasoning, and technical domains.

Model Pricing

Current Pricing

Feature Price (per 1M tokens)
Prompt $0.37
Completion $1.25
Input Cache Read $0.075

Price History

Available Endpoints
Provider Endpoint Name Context Length Pricing (Input) Pricing (Output)
Z.AI
Z.AI | z-ai/glm-5.3-flashx-20260918 1M $0.37 / 1M tokens $1.25 / 1M tokens
Benchmark Results
Benchmark Category Reasoning Strategy Free Executions Accuracy Cost Duration
Other Models by z-ai