Back to Models
Google: Gemini 3.5 Flash
Googletext9.1 / 10 Overall Rating
Google's Gemini 3.5 Flash is a multimodal model featuring a 1M-token context window and a 65,536-token max output ceiling. Designed for high-speed coding and parallel agentic execution, it provides lower latency than Pro-tier equivalents. At $1.50/1M input tokens, it offers lower API costs than GPT-4o and Claude 3.5 Sonnet.
Context Window
1049K
Knowledge Cutoff
2025
Max Output
16K
License
Proprietary
Where it Excels (Pros)
- 1M token context window
- 65K maximum token output
- Native audio and video parsing
- Cheaper than GPT-4o and Sonnet
Limitations (Cons)
- Pricier than 1.5 Flash models
- Text-only output capability
Benchmark Breakdown
Reasoning (MMLU)82%
Coding (HumanEval)78%
Mathematics (MATH)76%
Long-Context Retrieval95%
Pricing Matrix
Input cost / 1M tokens$1.50
Output cost / 1M tokens$9.00
Compare Google: Gemini 3.5 Flash with other frontier models: