Back to Models

StepFun: Step 3.7 Flash

Stepfuntext
8.7 / 10 Overall Rating

Step 3.7 Flash is a 196B-parameter Mixture-of-Experts multimodal model from StepFun activating roughly 11B parameters per token. It features native image and video understanding across a 262,144-token context window. At $0.20 per million input tokens, it delivers high-throughput multimodal processing at a fraction of GPT-4o's cost.

Context Window
262K
Knowledge Cutoff
2025
Max Output
16K
License
Proprietary

Where it Excels (Pros)

  • Extremely low API pricing
  • 262K token multimodal context
  • Native video and image support

Limitations (Cons)

  • Unproven complex coding capabilities
  • Lacks top-tier reasoning vs Sonnet

Benchmark Breakdown

Reasoning (MMLU)82%
Coding (HumanEval)78%
Mathematics (MATH)76%
Long-Context Retrieval95%

Pricing Matrix

Input cost / 1M tokens$0.20
Output cost / 1M tokens$1.15

Compare StepFun: Step 3.7 Flash with other frontier models:

Related Skills, Tools & Automations