Back to Models
NVIDIA: Nemotron 3.5 Lightning
Nvidiatext8.3 / 10 Overall Rating
NVIDIA Nemotron 3.5 Lightning is an open Mixture-of-Experts (MoE) model featuring 3B active out of 30B total parameters, optimized for high-throughput agentic execution. It delivers a massive 262K context window and an unprecedented 131K max output limit at a tiny fraction of frontier model pricing. It trades peak reasoning depth for extreme inference speed, high token volume, and low operational cost.
Context Window
262K
Knowledge Cutoff
2025
Max Output
16K
License
Proprietary
Where it Excels (Pros)
- 262K context with 131K output
- Extremely low API pricing
- High throughput for agentic workloads
- Efficient 3B active parameter MoE
Limitations (Cons)
- Lower reasoning depth than GPT-4o
- 3B active parameter complexity limits
Benchmark Breakdown
Reasoning (MMLU)82%
Coding (HumanEval)78%
Mathematics (MATH)76%
Long-Context Retrieval95%
Pricing Matrix
Input cost / 1M tokens$0.08
Output cost / 1M tokens$0.20
Compare NVIDIA: Nemotron 3.5 Lightning with other frontier models: