Back to Models
NVIDIA: Nemotron 3 Ultra (batch)
Nvidiatext9.0 / 10 Overall Rating
NVIDIA Nemotron 3 Ultra (Batch) is a 550B total parameter (55B active) hybrid Transformer-Mamba MoE model designed for complex reasoning and agent orchestration. It supports a 512K token context window and up to 461K output tokens per call via batch processing. Priced at $0.60 per million input tokens, it offers a cost-effective alternative to GPT-4o and Claude 3.5 Sonnet for bulk, long-context text processing.
Context Window
512K
Knowledge Cutoff
2025
Max Output
16K
License
Proprietary
Where it Excels (Pros)
- 512K context with 461K output limit
- Significantly lower pricing than GPT-4o
- Efficient hybrid Transformer-Mamba MoE architecture
Limitations (Cons)
- High latency from batch endpoint
- Text-only with no multimodal support
Benchmark Breakdown
Reasoning (MMLU)82%
Coding (HumanEval)78%
Mathematics (MATH)76%
Long-Context Retrieval95%
Pricing Matrix
Input cost / 1M tokens$0.60
Output cost / 1M tokens$3.60
Compare NVIDIA: Nemotron 3 Ultra (batch) with other frontier models: