Thinking Machines: Inkling Small

Text input Image input Audio input Text output
Author's Description

Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of...

Key Specifications
Cost
$$$$
Context
524K
Released
Jul 30, 2026
Speed
Ability
Reliability
Supported Parameters

This model supports the following parameters:

Top P Min P Logit Bias Frequency Penalty Include Reasoning Temperature Stop Presence Penalty Reasoning Max Tokens Seed
Features

This model supports the following features:

Reasoning
Performance Summary

Inkling Small, an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, demonstrates a balanced performance profile with notable strengths. It exhibits competitive response times, ranking in the 55th percentile across benchmarks, and offers moderate pricing, placing it in the 34th percentile. A standout feature is its exceptional reliability, boasting a 99% success rate, indicating consistent and usable responses. The model excels in several key areas. It achieves perfect accuracy in both General Knowledge and Ethics, with the former also being the most accurate model at its price point and speed. Its Coding and Mathematics capabilities are strong, scoring 94% and 95% accuracy respectively, placing it in the top quartiles for these categories. Reasoning is also a significant strength at 98% accuracy. While its Hallucinations accuracy is 96%, indicating a good ability to acknowledge uncertainty, its Instruction Following accuracy of 71% suggests room for improvement in handling complex, multi-layered directives. Email Classification is a moderate performer at 97% accuracy. Overall, Inkling Small presents itself as a reliable and capable model, particularly strong in knowledge-based tasks, coding, and ethical reasoning, with efficiency in mind.

Model Pricing

Current Pricing

Feature Price (per 1M tokens)
Prompt $0.58
Completion $1.44
Input Cache Read $0.116

Price History

Available Endpoints
Provider Endpoint Name Context Length Pricing (Input) Pricing (Output)
DeepInfra
DeepInfra | thinkingmachines/inkling-small-20260730 524K $0.58 / 1M tokens $1.44 / 1M tokens
Benchmark Results
Benchmark Category Reasoning Strategy Free Executions Accuracy Cost Duration
Other Models by thinkingmachines