Back to Models
StepFun: Step 3.7 Flash
Stepfuntext8.7 / 10 Overall Rating
Step 3.7 Flash is a 196B-parameter Mixture-of-Experts multimodal model from StepFun activating roughly 11B parameters per token. It features native image and video understanding across a 262,144-token context window. At $0.20 per million input tokens, it delivers high-throughput multimodal processing at a fraction of GPT-4o's cost.
Context Window
262K
Knowledge Cutoff
2025
Max Output
16K
License
Proprietary
Where it Excels (Pros)
- Extremely low API pricing
- 262K token multimodal context
- Native video and image support
Limitations (Cons)
- Unproven complex coding capabilities
- Lacks top-tier reasoning vs Sonnet
Benchmark Breakdown
Reasoning (MMLU)82%
Coding (HumanEval)78%
Mathematics (MATH)76%
Long-Context Retrieval95%
Pricing Matrix
Input cost / 1M tokens$0.20
Output cost / 1M tokens$1.15
Compare StepFun: Step 3.7 Flash with other frontier models: