Z.ai: GLM 5.3 Prime

Text input Text output
Author's Description

GLM-5.3-Prime is the high-speed variant of Z.ai's GLM-5.3, inheriting its full capabilities while delivering 1.5–2× the output throughput through inference acceleration. It supports text input and output with a 1M-token...

Key Specifications
Cost
$$$$$
Context
1M
Released
Sep 23, 2026
Speed
Ability
Reliability
Supported Parameters

This model supports the following parameters:

Max Tokens Frequency Penalty Temperature Tool Choice Include Reasoning Reasoning Top P Stop Seed Response Format Logprobs Tools Presence Penalty Top Logprobs
Features

This model supports the following features:

Tools Response Format Reasoning
Performance Summary

Z.ai's GLM-5.3-Prime, released on September 23, 2026, is positioned as a high-speed variant of the GLM-5.3 series, boasting 1.5–2x output throughput and a substantial 1M-token context length. In terms of speed, the model demonstrates competitive response times, ranking in the 53rd percentile across six benchmarks. However, its pricing tends to be at premium levels, placing it in the 12th percentile. A standout feature is its exceptional reliability, achieving a 100% success rate across all benchmarks, indicating minimal technical failures and consistent operability. Analyzing benchmark performance, GLM-5.3-Prime exhibits remarkable accuracy in General Knowledge and Ethics, achieving perfect 100% scores. For both these categories, it is highlighted as the most accurate model at its price point and among models of comparable speed. Its reasoning capabilities are also strong, with 98.0% accuracy, placing it in the 84th percentile. While its Hallucinations (98.0% accuracy) and Email Classification (98.0% accuracy) scores are solid, they are closer to the median. Instruction Following, at 73.0% accuracy, represents a relative area for potential improvement compared to its other high-performing categories. Overall, GLM-5.3-Prime's key strengths lie in its high reliability, perfect accuracy in knowledge and ethical reasoning, and strong general reasoning, making it a robust choice for demanding applications despite its premium cost.

Model Pricing

Current Pricing

Feature Price (per 1M tokens)
Prompt $2.8
Completion $8.8
Input Cache Read $0.56

Price History

Available Endpoints
Provider Endpoint Name Context Length Pricing (Input) Pricing (Output)
Alibaba
Alibaba | z-ai/glm-5.3-prime-20260921 1M $2.8 / 1M tokens $8.8 / 1M tokens
Benchmark Results
Benchmark Category Reasoning Strategy Free Executions Accuracy Cost Duration
Other Models by z-ai