ToolsThe story, in brief

Ling 3.0 Flash is now available on AI Gateway

Ant Group's Ling 3.0 Flash hits Vercel's AI Gateway. Free for 3 weeks. 124B MoE built for agentic workflows at production scale.

Paper-cut illustration of a coral software window opening into a three-dimensional drafting space.
New tools for building and creating with AI.AI illustration by KeyNews
The KeyNews take

Why it matters

Ant Group's latest open model is now accessible via Vercel's unified API layer, lowering friction for developers building multi-step agent applications. The free tier and token-efficient design (5.1B active params) signal competitive pressure on inference costs in the agentic inference market.

The key facts

14 to know
  1. Ling 3.0 Flash: 124B total parameters, 5.1B active per token (MoE architecture)

  2. 256K token context window

  3. Free access through August 3rd via AI Gateway

  4. Designed for agentic inference, coding agents, document work, long-context interactions

  5. Available via Vercel AI SDK with model identifier: inclusionai/ling-3.0-flash-free

  6. AI Gateway pricing: no markup, no platform fee on inference, includes BYOK support

  7. Thinking and non-thinking modes supported

  8. Ling 3.0 Flash: 124B total parameters, ~5.1B active per token

  9. Mixture-of-Experts architecture optimized for token efficiency

  10. Free access through August 3rd, 2026

  11. Available via Vercel AI Gateway with no platform markup or inference fees

  12. Targeting high-frequency agentic workflows, coding agents, document work, long-context multi-turn interactions

  13. Supports thinking and non-thinking modes

  14. AI Gateway includes custom reporting, Zero Data Retention, budgets for API keys, routing rules

Go to the source

Vercel Blogvercel.com

Publisher excerpt: from Ant Group is now available on AI Gateway.Ling 3.0 Flash The model is free to use for the next three weeks, through August 3rd. Ling 3.0 Flash is a Mixture-of-Experts model with 124B total parameters and about 5.1B active per token. It has a 256K token context window and runs in thinking and…
Read original report
Back to today's editionMore tools news

Keep reading

Related stories

More from Tools