Back to Models
Thinking Machines: Inkling (batch)
Thinkingmachinestext9.1 / 10 Overall Rating
Inkling (batch) is an open-weight, 975B-parameter multimodal Mixture-of-Experts model (41B active) developed by Thinking Machines Lab. It accepts text, image, and audio inputs across a 524K context window, offering high-volume reasoning and coding at a fraction of GPT-4o and Claude 3.5 Sonnet API costs.
Context Window
524K
Knowledge Cutoff
2025
Max Output
16K
License
Proprietary
Where it Excels (Pros)
- 524K token context window
- Lower API cost than GPT-4o
- 471K max output tokens
- Native audio and image inputs
Limitations (Cons)
- Batch API introduces execution latency
- Requires massive self-hosting infrastructure
Benchmark Breakdown
Reasoning (MMLU)82%
Coding (HumanEval)78%
Mathematics (MATH)76%
Long-Context Retrieval95%
Pricing Matrix
Input cost / 1M tokens$1.00
Output cost / 1M tokens$4.05
Compare Thinking Machines: Inkling (batch) with other frontier models: