NVIDIA: Nemotron 3.5 Lightning (free)

Text input Text output
Author's Description

NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...

Key Specifications
Cost
$$$
Context
262K
Parameters
30B (Rumoured)
Released
Aug 11, 2026
Speed
Ability
Reliability
Supported Parameters

This model supports the following parameters:

Reasoning Include Reasoning Frequency Penalty Top P Presence Penalty Stop Logprobs Max Tokens Temperature Top Logprobs Seed Response Format Structured Outputs
Features

This model supports the following features:

Structured Outputs Response Format Reasoning
Performance Summary

NVIDIA Nemotron 3.5 Lightning demonstrates moderate speed performance, ranking in the 27th percentile across benchmarks. It offers cost-effective solutions, placing in the 69th percentile for pricing. A significant strength is its strong reliability, achieving a 94% success rate, indicating consistent and usable responses. In terms of specific performance, the model excels in hallucination prevention, achieving perfect 100% accuracy, making it the most accurate model at its price point and speed for this task. It also shows strong performance in Email Classification (96.0% accuracy) and General Knowledge (88.8% accuracy), though its General Knowledge accuracy is in the lower quartile. The model performs reasonably well in Mathematics (88.0%) and Coding (81.8%). However, it exhibits notable weaknesses in Reasoning (57.1% accuracy) and Instruction Following (70.7% accuracy), suggesting areas for improvement in complex logical deduction and multi-step command execution. Its Ethics performance, while 88.0% accurate, ranks in the lower 19th percentile, indicating it may not be as strong as other models in this domain. Overall, Nemotron 3.5 Lightning is a reliable and cost-efficient option, particularly strong in avoiding hallucinations, but has room for growth in complex reasoning and instruction adherence.

Model Pricing

Current Pricing

Feature Price (per 1M tokens)
Prompt $0.1
Completion $0.25
Input Cache Read $0.05

Price History

Available Endpoints
Provider Endpoint Name Context Length Pricing (Input) Pricing (Output)
CoreWeave
CoreWeave | nvidia/nemotron-3.5-lightning-20260807 262K $0.1 / 1M tokens $0.25 / 1M tokens
DeepInfra
DeepInfra | nvidia/nemotron-3.5-lightning-20260807 28K $0.05 / 1M tokens $0.2 / 1M tokens
Benchmark Results
Benchmark Category Reasoning Strategy Free Executions Accuracy Cost Duration
Other Models by nvidia