Z.ai: GLM 5.3 Flash

Text input Image input Video input Text output
Author's Description

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...

Key Specifications
Cost
$$$
Context
1M
Released
Aug 26, 2026
Speed
Ability
Reliability
Supported Parameters

This model supports the following parameters:

Temperature Tool Choice Response Format Include Reasoning Reasoning Top P Tools Max Tokens
Features

This model supports the following features:

Reasoning Tools Response Format
Performance Summary

Z.ai's GLM 5.3 Flash, created on August 26, 2026, is a native multimodal model designed for efficient coding and long-horizon agent tasks, leveraging a hybrid sparse and linear attention architecture for accurate long-context behavior. The model demonstrates moderate speed performance, ranking in the 34th percentile across seven benchmarks. It generally offers cost-effective solutions, placing in the 69th percentile for price. A standout feature is its exceptional reliability, boasting a 99% success rate, indicating consistent and usable responses. In terms of benchmark performance, GLM 5.3 Flash exhibits perfect accuracy in Hallucinations (100.0%) and General Knowledge (100.0%), positioning it as a top performer in these areas, often being the most accurate model at its price point and speed. It also shows strong capabilities in Reasoning (98.0% accuracy, 91st percentile) and Coding (95.0% accuracy, 90th percentile), making it well-suited for its intended applications. While its Instruction Following is solid at 77.0% accuracy (79th percentile), its Ethics performance is comparatively lower at 96.0% accuracy (25th percentile), suggesting a potential area for improvement despite still being a high score. Email Classification is strong at 98.0% accuracy. Overall, its key strengths lie in its accuracy for knowledge-based and reasoning tasks, coupled with high reliability and cost-effectiveness.

Model Pricing

Current Pricing

Feature Price (per 1M tokens)
Prompt $0.075
Completion $0.25
Input Cache Read $0.015

Price History

Available Endpoints
Provider Endpoint Name Context Length Pricing (Input) Pricing (Output)
Z.AI
Z.AI | z-ai/glm-5.3-flash-20260826 1M $0.075 / 1M tokens $0.25 / 1M tokens
Benchmark Results
Benchmark Category Reasoning Strategy Free Executions Accuracy Cost Duration
Other Models by z-ai