Back to Models

Thinking Machines: Inkling (batch)

Thinkingmachinestext
9.1 / 10 Overall Rating

Inkling (batch) is an open-weight, 975B-parameter multimodal Mixture-of-Experts model (41B active) developed by Thinking Machines Lab. It accepts text, image, and audio inputs across a 524K context window, offering high-volume reasoning and coding at a fraction of GPT-4o and Claude 3.5 Sonnet API costs.

Context Window
524K
Knowledge Cutoff
2025
Max Output
16K
License
Proprietary

Where it Excels (Pros)

  • 524K token context window
  • Lower API cost than GPT-4o
  • 471K max output tokens
  • Native audio and image inputs

Limitations (Cons)

  • Batch API introduces execution latency
  • Requires massive self-hosting infrastructure

Benchmark Breakdown

Reasoning (MMLU)82%
Coding (HumanEval)78%
Mathematics (MATH)76%
Long-Context Retrieval95%

Pricing Matrix

Input cost / 1M tokens$1.00
Output cost / 1M tokens$4.05

Compare Thinking Machines: Inkling (batch) with other frontier models:

Related Skills, Tools & Automations