FrontierThe story, in brief

StepFun Releases Step 3.7 Flash: A 198B MoE Vision-Language Model for Coding Agents and Search Workflows

198B parameters. Native vision. 256k context. StepFun's Step 3.7 Flash is built for agents—not chat.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

StepFun enters the vision-language model competition with a MoE architecture optimized for agentic workflows and coding tasks, signaling intensifying competition in multimodal reasoning models designed for enterprise automation.

The key facts

6 to know
  1. Step 3.7 Flash: 198B MoE parameters

  2. Native vision-language capability

  3. 256k context window

  4. Advisor Mode feature

  5. Positioned for coding agents and search workflows

  6. Published May 29, 2026

Go to the source

MarkTechPostmarktechpost.com

Publisher excerpt: StepFun releases Step 3.7 Flash, a 198B MoE model with native vision, 256k context, and Advisor Mode.
Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier