Back to Models
Thinking Machines: Inkling
Thinkingmachinestext9.1 / 10 Overall Rating
Developed by Thinking Machines Lab, Inkling is an open-weight mixture-of-experts model with 41B active parameters out of 975B total. It features a 1M-token context window and handles text, image, and audio inputs for reasoning, coding, and agentic tasks. At $0.95/1M input tokens and a 262K output limit, it offers deep context processing at a fraction of GPT-4o costs.
Context Window
1049K
Knowledge Cutoff
2025
Max Output
16K
License
Proprietary
Where it Excels (Pros)
- 1M token input context window
- Massive 262K token output limit
- Cheaper than GPT-4o and Claude
- Open-weight architecture with 975B parameters
Limitations (Cons)
- High VRAM requirements for self-hosting
- Massive total parameter storage footprint
Benchmark Breakdown
Reasoning (MMLU)82%
Coding (HumanEval)78%
Mathematics (MATH)76%
Long-Context Retrieval95%
Pricing Matrix
Input cost / 1M tokens$0.95
Output cost / 1M tokens$4.05
Compare Thinking Machines: Inkling with other frontier models: