Back to Models

Thinking Machines: Inkling

Thinkingmachinestext
9.1 / 10 Overall Rating

Developed by Thinking Machines Lab, Inkling is an open-weight mixture-of-experts model with 41B active parameters out of 975B total. It features a 1M-token context window and handles text, image, and audio inputs for reasoning, coding, and agentic tasks. At $0.95/1M input tokens and a 262K output limit, it offers deep context processing at a fraction of GPT-4o costs.

Context Window
1049K
Knowledge Cutoff
2025
Max Output
16K
License
Proprietary

Where it Excels (Pros)

  • 1M token input context window
  • Massive 262K token output limit
  • Cheaper than GPT-4o and Claude
  • Open-weight architecture with 975B parameters

Limitations (Cons)

  • High VRAM requirements for self-hosting
  • Massive total parameter storage footprint

Benchmark Breakdown

Reasoning (MMLU)82%
Coding (HumanEval)78%
Mathematics (MATH)76%
Long-Context Retrieval95%

Pricing Matrix

Input cost / 1M tokens$0.95
Output cost / 1M tokens$4.05

Compare Thinking Machines: Inkling with other frontier models:

Related Skills, Tools & Automations