Back to Models

Google: Gemini 3.6 Flash (batch)

Googletext
8.9 / 10 Overall Rating

Google's Gemini 3.6 Flash (Batch) provides asynchronous inference with a 1,048,576 token context window and 65,536 max output tokens. Designed for high-volume background processing, it handles coding, agentic workflows, and long-document analysis at lower batch pricing. It offers extreme context capacity at a fraction of GPT-4o and Claude 3.5 Sonnet API costs.

Context Window
1049K
Knowledge Cutoff
2025
Max Output
16K
License
Proprietary

Where it Excels (Pros)

  • 1M token input context window
  • 65K max output token capacity
  • Extremely low batch API pricing
  • Native video and audio processing

Limitations (Cons)

  • High latency batch endpoint only
  • Unsuitable for real-time interactive apps
  • Weaker logic reasoning than Sonnet

Benchmark Breakdown

Reasoning (MMLU)82%
Coding (HumanEval)78%
Mathematics (MATH)76%
Long-Context Retrieval95%

Pricing Matrix

Input cost / 1M tokens$0.38
Output cost / 1M tokens$1.88

Compare Google: Gemini 3.6 Flash (batch) with other frontier models:

Related Skills, Tools & Automations