Back to Models
Ling-3.0-flash
Inclusionaitext8.4 / 10 Overall Rating
Ling-3.0-flash is a 124B-parameter Mixture-of-Experts (MoE) text model developed by inclusionai, activating 5.1B parameters per token. It features a 262K token context window, a 32K token max output, and extremely low API costs ($0.02/$0.06 per 1M tokens). The architecture targets high-throughput agentic workflows and budget-focused enterprise text tasks.
Context Window
262K
Knowledge Cutoff
2025
Max Output
16K
License
Proprietary
Where it Excels (Pros)
- Extremely low API inference pricing
- 262K context window beats GPT-4o
- High 32K max output tokens
- Efficient 5.1B active parameter MoE
Limitations (Cons)
- Weaker complex reasoning than Claude 3.5
- Limited public benchmark verification data
Benchmark Breakdown
Reasoning (MMLU)82%
Coding (HumanEval)78%
Mathematics (MATH)76%
Long-Context Retrieval95%
Pricing Matrix
Input cost / 1M tokens$0.02
Output cost / 1M tokens$0.06
Compare Ling-3.0-flash with other frontier models: