Back to Models

Google: Gemini 3.5 Flash

Googletext
9.1 / 10 Overall Rating

Google's Gemini 3.5 Flash is a multimodal model featuring a 1M-token context window and a 65,536-token max output ceiling. Designed for high-speed coding and parallel agentic execution, it provides lower latency than Pro-tier equivalents. At $1.50/1M input tokens, it offers lower API costs than GPT-4o and Claude 3.5 Sonnet.

Context Window
1049K
Knowledge Cutoff
2025
Max Output
16K
License
Proprietary

Where it Excels (Pros)

  • 1M token context window
  • 65K maximum token output
  • Native audio and video parsing
  • Cheaper than GPT-4o and Sonnet

Limitations (Cons)

  • Pricier than 1.5 Flash models
  • Text-only output capability

Benchmark Breakdown

Reasoning (MMLU)82%
Coding (HumanEval)78%
Mathematics (MATH)76%
Long-Context Retrieval95%

Pricing Matrix

Input cost / 1M tokens$1.50
Output cost / 1M tokens$9.00

Compare Google: Gemini 3.5 Flash with other frontier models:

Related Skills, Tools & Automations