Author's Description
NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...
Key Specifications
Supported Parameters
This model supports the following parameters:
Features
This model supports the following features:
Performance Summary
NVIDIA Nemotron 3.5 Lightning demonstrates moderate speed performance, ranking in the 27th percentile across benchmarks. It offers cost-effective solutions, placing in the 69th percentile for pricing. A significant strength is its strong reliability, achieving a 94% success rate, indicating consistent and usable responses. In terms of specific performance, the model excels in hallucination prevention, achieving perfect 100% accuracy, making it the most accurate model at its price point and speed for this task. It also shows strong performance in Email Classification (96.0% accuracy) and General Knowledge (88.8% accuracy), though its General Knowledge accuracy is in the lower quartile. The model performs reasonably well in Mathematics (88.0%) and Coding (81.8%). However, it exhibits notable weaknesses in Reasoning (57.1% accuracy) and Instruction Following (70.7% accuracy), suggesting areas for improvement in complex logical deduction and multi-step command execution. Its Ethics performance, while 88.0% accurate, ranks in the lower 19th percentile, indicating it may not be as strong as other models in this domain. Overall, Nemotron 3.5 Lightning is a reliable and cost-efficient option, particularly strong in avoiding hallucinations, but has room for growth in complex reasoning and instruction adherence.
Model Pricing
Current Pricing
| Feature | Price (per 1M tokens) |
|---|---|
| Prompt | $0.1 |
| Completion | $0.25 |
| Input Cache Read | $0.05 |
Price History
Available Endpoints
| Provider | Endpoint Name | Context Length | Pricing (Input) | Pricing (Output) |
|---|---|---|---|---|
|
CoreWeave
|
CoreWeave | nvidia/nemotron-3.5-lightning-20260807 | 262K | $0.1 / 1M tokens | $0.25 / 1M tokens |
|
DeepInfra
|
DeepInfra | nvidia/nemotron-3.5-lightning-20260807 | 28K | $0.05 / 1M tokens | $0.2 / 1M tokens |
Benchmark Results
| Benchmark | Category | Reasoning | Strategy | Free | Executions | Accuracy | Cost | Duration |
|---|
Other Models by nvidia
|
|
Released | Params | Context |
|
Speed | Ability | Cost |
|---|---|---|---|---|---|---|---|
| NVIDIA: Nemotron 3.5 Content Safety (free) Unavailable | Jun 04, 2026 | ~4B | N/A |
Text input
Image input
Text output
|
— | — | — |
| NVIDIA: Nemotron 3 Ultra (free) | Jun 03, 2026 | 550B | 512K |
Text input
Text output
|
★ | ★ | $$$$ |
| NVIDIA: Nemotron 3 Nano Omni (free) Unavailable | Apr 28, 2026 | 30B | N/A |
Text input
Video input
Image input
Audio input
Text output
|
— | — | — |
| NVIDIA: Nemotron 3 Super (free) | Mar 11, 2026 | 120B | 262K |
Text input
Text output
|
★★★ | ★★★ | $$$$ |
| NVIDIA: Nemotron 3 Nano 30B A3B (free) | Dec 14, 2025 | 30B | 262K |
Text input
Text output
|
★★★ | ★★★★★ | $$$ |
| NVIDIA: Nemotron Nano 12B 2 VL (free) | Oct 28, 2025 | 12B | 131K |
Text input
Video input
Image input
Text output
|
★ | ★★ | $$$$ |
| NVIDIA: Llama 3.3 Nemotron Super 49B V1.5 Unavailable | Oct 10, 2025 | 49B | 131K |
Text input
Text output
|
★★ | ★★★★ | $$$$ |
| NVIDIA: Nemotron Nano 9B V2 (free) | Sep 05, 2025 | 9B | 128K |
Text input
Text output
|
★ | ★★ | $ |
| NVIDIA: Llama 3.3 Nemotron Super 49B v1 Unavailable | Apr 08, 2025 | 49B | 131K |
Text input
Text output
|
★★★★ | ★★ | $$ |
| NVIDIA: Llama 3.1 Nemotron Ultra 253B v1 Unavailable | Apr 08, 2025 | 253B | 131K |
Text input
Text output
|
★★ | ★★ | $$$$ |
| NVIDIA: Llama 3.1 Nemotron 70B Instruct Unavailable | Oct 14, 2024 | 70B | 131K |
Text input
Text output
|
★★★ | ★★ | $$ |