Back to Models
Google: Gemini 3.5 Flash Lite (batch)
Googletext8.3 / 10 Overall Rating
Google Gemini 3.5 Flash Lite (batch) is a lightweight, multimodal model engineered for high-volume asynchronous tasks and subagent workflows. It offers a 1-million-token context window and a 65,536-token output limit at a fraction of the cost of standard frontier models. It sacrifices real-time response latency and deep reasoning capabilities to achieve high cost-efficiency.
Context Window
1049K
Knowledge Cutoff
2025
Max Output
16K
License
Proprietary
Where it Excels (Pros)
- 1M token context window
- Extremely low input output cost
- Native audio and video multimodal support
Limitations (Cons)
- Lower reasoning depth than GPT-4o
- Batch execution introduces high latency
Benchmark Breakdown
Reasoning (MMLU)82%
Coding (HumanEval)78%
Mathematics (MATH)76%
Long-Context Retrieval95%
Pricing Matrix
Input cost / 1M tokens$0.15
Output cost / 1M tokens$1.25
Compare Google: Gemini 3.5 Flash Lite (batch) with other frontier models: