StepFun Releases Step 3.7 Flash: A 198B MoE Vision-Language Model for Coding Agents and Search Workflows
198B parameters. Native vision. 256k context. StepFun's Step 3.7 Flash is built for agents—not chat.

Why it matters
StepFun enters the vision-language model competition with a MoE architecture optimized for agentic workflows and coding tasks, signaling intensifying competition in multimodal reasoning models designed for enterprise automation.
The key facts
6 to knowStep 3.7 Flash: 198B MoE parameters
Native vision-language capability
256k context window
Advisor Mode feature
Positioned for coding agents and search workflows
Published May 29, 2026
Go to the source
MarkTechPostmarktechpost.com
Publisher excerpt: StepFun releases Step 3.7 Flash, a 198B MoE model with native vision, 256k context, and Advisor Mode.