Back to Models
Google: Gemini 3.6 Flash (batch)
Googletext8.9 / 10 Overall Rating
Google's Gemini 3.6 Flash (Batch) provides asynchronous inference with a 1,048,576 token context window and 65,536 max output tokens. Designed for high-volume background processing, it handles coding, agentic workflows, and long-document analysis at lower batch pricing. It offers extreme context capacity at a fraction of GPT-4o and Claude 3.5 Sonnet API costs.
Context Window
1049K
Knowledge Cutoff
2025
Max Output
16K
License
Proprietary
Where it Excels (Pros)
- 1M token input context window
- 65K max output token capacity
- Extremely low batch API pricing
- Native video and audio processing
Limitations (Cons)
- High latency batch endpoint only
- Unsuitable for real-time interactive apps
- Weaker logic reasoning than Sonnet
Benchmark Breakdown
Reasoning (MMLU)82%
Coding (HumanEval)78%
Mathematics (MATH)76%
Long-Context Retrieval95%
Pricing Matrix
Input cost / 1M tokens$0.38
Output cost / 1M tokens$1.88
Compare Google: Gemini 3.6 Flash (batch) with other frontier models: