Qwen: Qwen3.8 2.4T A95B (batch)

Text input Text output
Author's Description

Qwen3.8 2.4T A95B is an open-weight sparse mixture-of-experts model from Qwen and the open-weight variant of [Qwen3.8 Max](/qwen/qwen3.8-max), with 95 billion active parameters out of 2.4 trillion total. It is...

Key Specifications
Cost
$$$$$
Context
262K
Parameters
95B
Released
Aug 12, 2026
Speed
Ability
Reliability
Supported Parameters

This model supports the following parameters:

Top P Max Tokens Temperature Tool Choice Response Format Logprobs Include Reasoning Tools Structured Outputs Top Logprobs Reasoning
Features

This model supports the following features:

Tools Response Format Structured Outputs Reasoning
Performance Summary

Qwen3.8 2.4T A95B, an open-weight sparse mixture-of-experts model, demonstrates exceptional performance across several key metrics. It consistently ranks among the fastest models available and offers highly competitive pricing, making it an attractive option for cost-sensitive applications. The model also exhibits outstanding reliability, with a 98% success rate across benchmarks, indicating minimal technical failures. In terms of specific capabilities, Qwen3.8 2.4T A95B achieves perfect accuracy in both General Knowledge and Ethics benchmarks, often being the most accurate model at its price point and speed. It shows strong performance in Coding (85th percentile), Reasoning (73rd percentile), and Mathematics (80th percentile). Email Classification is also robust at 99.0% accuracy. A notable weakness is its Instruction Following, where it scored 0.0% accuracy, suggesting this area requires significant improvement. Its Hallucinations score of 85.1% accuracy, while not the lowest, indicates room for improvement in acknowledging uncertainty. Overall, the model presents a compelling balance of speed, cost-effectiveness, and high accuracy in critical domains, despite specific areas needing development.

Model Pricing

Current Pricing

Feature Price (per 1M tokens)
Prompt $2
Completion $6
Input Cache Read $0.25

Price History

Available Endpoints
Provider Endpoint Name Context Length Pricing (Input) Pricing (Output)
DigitalOcean
DigitalOcean | qwen/qwen3.8-2.4t-a95b-20260812 262K $2 / 1M tokens $6 / 1M tokens
Fireworks
Fireworks | qwen/qwen3.8-2.4t-a95b-20260812 262K $2 / 1M tokens $6 / 1M tokens
DeepInfra
DeepInfra | qwen/qwen3.8-2.4t-a95b-20260812 262K $2 / 1M tokens $6 / 1M tokens
Modal
Modal | qwen/qwen3.8-2.4t-a95b-20260812 1M $2 / 1M tokens $6 / 1M tokens
Together
Together | qwen/qwen3.8-2.4t-a95b-20260812 1M $2 / 1M tokens $6 / 1M tokens
Alibaba
Alibaba | qwen/qwen3.8-2.4t-a95b-20260812 1M $2 / 1M tokens $6 / 1M tokens
SiliconFlow
SiliconFlow | qwen/qwen3.8-2.4t-a95b-20260812 1M $2 / 1M tokens $6 / 1M tokens
Venice
Venice | qwen/qwen3.8-2.4t-a95b-20260812 262K $2 / 1M tokens $6 / 1M tokens
Novita
Novita | qwen/qwen3.8-2.4t-a95b-20260812 1M $2 / 1M tokens $6 / 1M tokens
Modal
Modal | qwen/qwen3.8-2.4t-a95b-20260812 1M $2 / 1M tokens $6 / 1M tokens
Benchmark Results
Benchmark Category Reasoning Strategy Free Executions Accuracy Cost Duration
Other Models by qwen