Back to Models

NVIDIA: Nemotron 3 Ultra (batch)

Nvidiatext
9.0 / 10 Overall Rating

NVIDIA Nemotron 3 Ultra (Batch) is a 550B total parameter (55B active) hybrid Transformer-Mamba MoE model designed for complex reasoning and agent orchestration. It supports a 512K token context window and up to 461K output tokens per call via batch processing. Priced at $0.60 per million input tokens, it offers a cost-effective alternative to GPT-4o and Claude 3.5 Sonnet for bulk, long-context text processing.

Context Window
512K
Knowledge Cutoff
2025
Max Output
16K
License
Proprietary

Where it Excels (Pros)

  • 512K context with 461K output limit
  • Significantly lower pricing than GPT-4o
  • Efficient hybrid Transformer-Mamba MoE architecture

Limitations (Cons)

  • High latency from batch endpoint
  • Text-only with no multimodal support

Benchmark Breakdown

Reasoning (MMLU)82%
Coding (HumanEval)78%
Mathematics (MATH)76%
Long-Context Retrieval95%

Pricing Matrix

Input cost / 1M tokens$0.60
Output cost / 1M tokens$3.60

Compare NVIDIA: Nemotron 3 Ultra (batch) with other frontier models:

Related Skills, Tools & Automations