Z.ai: GLM 5.3 Flash (batch)

Video input Image input Text input Text output
Author's Description

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...

Key Specifications
Cost
$$$
Context
1M
Released
Aug 26, 2026
Speed
Ability
Reliability
Supported Parameters

This model supports the following parameters:

Top P Max Tokens Temperature Tool Choice Response Format Include Reasoning Tools Reasoning
Features

This model supports the following features:

Tools Response Format Reasoning
Performance Summary

Z.ai's GLM 5.3 Flash, a native multimodal model released on August 26, 2026, demonstrates a balanced performance profile, particularly excelling in accuracy and reliability. While its speed performance is moderate, ranking in the 32nd percentile across benchmarks, it offers cost-effective solutions, typically falling within the 67th percentile for pricing. A standout feature is its exceptional reliability, boasting a 99% success rate, indicating minimal technical failures and consistent response generation. The model exhibits perfect accuracy in Hallucinations and General Knowledge benchmarks, showcasing its ability to acknowledge uncertainty and its broad factual understanding. It also performs strongly in Coding (95% accuracy, 90th percentile) and Reasoning (98% accuracy, 91st percentile), aligning with its description for efficient coding and long-horizon agent tasks. Instruction Following and Mathematics also show robust performance at 77% and 95% accuracy respectively. While Email Classification is solid at 98%, its Ethics score of 96% places it in the 25th percentile, suggesting this area could be a relative weakness compared to its other high-performing categories. Overall, GLM 5.3 Flash is a highly reliable and accurate model, particularly strong in knowledge, reasoning, and coding, with competitive pricing.

Model Pricing

Current Pricing

Feature Price (per 1M tokens)
Prompt $0.15
Completion $0.5
Input Cache Read $0.03

Price History

Available Endpoints
Provider Endpoint Name Context Length Pricing (Input) Pricing (Output)
Z.AI
Z.AI | z-ai/glm-5.3-flash-20260826 1M $0.15 / 1M tokens $0.5 / 1M tokens
Novita
Novita | z-ai/glm-5.3-flash-20260826 1M $0.132 / 1M tokens $0.44 / 1M tokens
Cloudflare
Cloudflare | z-ai/glm-5.3-flash-20260826 1M $0.15 / 1M tokens $0.5 / 1M tokens
GMICloud
GMICloud | z-ai/glm-5.3-flash-20260826 1M $0.105 / 1M tokens $0.35 / 1M tokens
DeepInfra
DeepInfra | z-ai/glm-5.3-flash-20260826 1M $0.075 / 1M tokens $0.25 / 1M tokens
Io Net
Io Net | z-ai/glm-5.3-flash-20260826 262K $0.15 / 1M tokens $0.5 / 1M tokens
BaseTen
BaseTen | z-ai/glm-5.3-flash-20260826 1M $0.15 / 1M tokens $0.5 / 1M tokens
Modal
Modal | z-ai/glm-5.3-flash-20260826 1M $0.45 / 1M tokens $1.5 / 1M tokens
Venice
Venice | z-ai/glm-5.3-flash-20260826 1M $0.15 / 1M tokens $0.5 / 1M tokens
Parasail
Parasail | z-ai/glm-5.3-flash-20260826 1M $0.15 / 1M tokens $0.5 / 1M tokens
Reka
Reka | z-ai/glm-5.3-flash-20260826 262K $0.225 / 1M tokens $0.75 / 1M tokens
Together
Together | z-ai/glm-5.3-flash-20260826 1M $0.15 / 1M tokens $0.5 / 1M tokens
Wafer
Wafer | z-ai/glm-5.3-flash-20260826 1M $0.1 / 1M tokens $0.35 / 1M tokens
Morph
Morph | z-ai/glm-5.3-flash-20260826 1M $0.1 / 1M tokens $0.35 / 1M tokens
DigitalOcean
DigitalOcean | z-ai/glm-5.3-flash-20260826 1M $0.15 / 1M tokens $0.5 / 1M tokens
Friendli
Friendli | z-ai/glm-5.3-flash-20260826 1M $0.15 / 1M tokens $0.5 / 1M tokens
Phala
Phala | z-ai/glm-5.3-flash-20260826 1M $0.15 / 1M tokens $0.5 / 1M tokens
Makora
Makora | z-ai/glm-5.3-flash-20260826 1M $0.075 / 1M tokens $0.25 / 1M tokens
Relace
Relace | z-ai/glm-5.3-flash-20260826 1M $0.075 / 1M tokens $0.25 / 1M tokens
SiliconFlow
SiliconFlow | z-ai/glm-5.3-flash-20260826 1M $0.15 / 1M tokens $0.5 / 1M tokens
Fireworks
Fireworks | z-ai/glm-5.3-flash-20260826 1M $0.15 / 1M tokens $0.5 / 1M tokens
Relace
Relace | z-ai/glm-5.3-flash-20260826 1M $0.075 / 1M tokens $0.25 / 1M tokens
StreamLake
StreamLake | z-ai/glm-5.3-flash-20260826 1M $0.075 / 1M tokens $0.25 / 1M tokens
NextBit
NextBit | z-ai/glm-5.3-flash-20260826 1M $0.2 / 1M tokens $0.675 / 1M tokens
Sail Research
Sail Research | z-ai/glm-5.3-flash-20260826 1M $0.15 / 1M tokens $0.5 / 1M tokens
DeepInfra
DeepInfra | z-ai/glm-5.3-flash-20260826 1M $0.075 / 1M tokens $0.25 / 1M tokens
CoreWeave
CoreWeave | z-ai/glm-5.3-flash-20260826 1M $0.075 / 1M tokens $0.25 / 1M tokens
StreamLake
StreamLake | z-ai/glm-5.3-flash-20260826 1M $0.105 / 1M tokens $0.349 / 1M tokens
Crusoe
Crusoe | z-ai/glm-5.3-flash-20260826 1M $0.15 / 1M tokens $0.5 / 1M tokens
Relace
Relace | z-ai/glm-5.3-flash-20260826 1M $0.09 / 1M tokens $0.3 / 1M tokens
Near AI
Near AI | z-ai/glm-5.3-flash-20260826 1M $0.075 / 1M tokens $0.25 / 1M tokens
OpenInference
OpenInference | z-ai/glm-5.3-flash-20260826 1M $0.075 / 1M tokens $0.25 / 1M tokens
AtlasCloud
AtlasCloud | z-ai/glm-5.3-flash-20260826 1M $0.15 / 1M tokens $0.5 / 1M tokens
CoreWeave
CoreWeave | z-ai/glm-5.3-flash-20260826 1M $0.15 / 1M tokens $0.5 / 1M tokens
Inceptron
Inceptron | z-ai/glm-5.3-flash-20260826 1M $0.15 / 1M tokens $0.5 / 1M tokens
InferenceNet
InferenceNet | z-ai/glm-5.3-flash-20260826 1M $0.09 / 1M tokens $0.28 / 1M tokens
Benchmark Results
Benchmark Category Reasoning Strategy Free Executions Accuracy Cost Duration
Other Models by z-ai