Back to Models

Ling-3.0-flash

Inclusionaitext
8.4 / 10 Overall Rating

Ling-3.0-flash is a 124B-parameter Mixture-of-Experts (MoE) text model developed by inclusionai, activating 5.1B parameters per token. It features a 262K token context window, a 32K token max output, and extremely low API costs ($0.02/$0.06 per 1M tokens). The architecture targets high-throughput agentic workflows and budget-focused enterprise text tasks.

Context Window
262K
Knowledge Cutoff
2025
Max Output
16K
License
Proprietary

Where it Excels (Pros)

  • Extremely low API inference pricing
  • 262K context window beats GPT-4o
  • High 32K max output tokens
  • Efficient 5.1B active parameter MoE

Limitations (Cons)

  • Weaker complex reasoning than Claude 3.5
  • Limited public benchmark verification data

Benchmark Breakdown

Reasoning (MMLU)82%
Coding (HumanEval)78%
Mathematics (MATH)76%
Long-Context Retrieval95%

Pricing Matrix

Input cost / 1M tokens$0.02
Output cost / 1M tokens$0.06

Compare Ling-3.0-flash with other frontier models:

Related Skills, Tools & Automations