NVIDIA: Llama 3.3 Nemotron Super 49B V1.5

Text input Text output Unavailable
Author's Description

Llama-3.3-Nemotron-Super-49B-v1.5 is a 49B-parameter, English-centric reasoning/chat model derived from Meta’s Llama-3.3-70B-Instruct with a 128K context. It’s post-trained for agentic workflows (RAG, tool calling) via SFT across math, code, science, and...

Key Specifications
Cost
$$$$
Context
131K
Parameters
49B
Released
Oct 10, 2025
Speed
Ability
Reliability
Supported Parameters

This model supports the following parameters:

Stop Max Tokens Seed Reasoning Top P Frequency Penalty Presence Penalty Temperature Include Reasoning Tools Tool Choice Logit Bias Response Format Min P
Features

This model supports the following features:

Response Format Tools Reasoning
Performance Summary

NVIDIA's Llama 3.3 Nemotron Super 49B V1.5 demonstrates moderate speed performance, ranking in the 22nd percentile across various benchmarks. It offers competitive pricing, placing in the 43rd percentile. The model exhibits exceptional reliability with a 97% success rate, indicating minimal technical failures and consistent response generation. In terms of benchmark performance, the model shows strong capabilities in General Knowledge (99.0% accuracy), Email Classification (99.0% accuracy), and Reasoning (92.0% accuracy). Its Ethics performance is perfect at 100.0% accuracy, making it the most accurate model at its price point and among models of similar speed. Coding accuracy is solid at 89.0%, and it achieves 85.0% in Mathematics. A notable weakness is its Hallucinations score of 90.0% accuracy, which, while decent, places it in the 40th percentile, suggesting room for improvement in acknowledging uncertainty. Instruction Following is moderate at 60.2%. Overall, the model is well-suited for agentic workflows, balancing accuracy and cost-efficiency, particularly in areas requiring strong reasoning and reliable tool use.

Model Pricing

Current Pricing

Feature Price (per 1M tokens)
Prompt $0.4
Completion $0.4

Price History

Available Endpoints
Provider Endpoint Name Context Length Pricing (Input) Pricing (Output)
DeepInfra
DeepInfra | nvidia/llama-3.3-nemotron-super-49b-v1.5 131K $0.4 / 1M tokens $0.4 / 1M tokens
Benchmark Results
Benchmark Category Reasoning Strategy Free Executions Accuracy Cost Duration
Other Models by nvidia