Ling 3.0 Flash is now available on AI Gateway
Ant Group's Ling 3.0 Flash hits Vercel's AI Gateway. Free for 3 weeks. 124B MoE built for agentic workflows at production scale.

Why it matters
Ant Group's latest open model is now accessible via Vercel's unified API layer, lowering friction for developers building multi-step agent applications. The free tier and token-efficient design (5.1B active params) signal competitive pressure on inference costs in the agentic inference market.
The key facts
14 to knowLing 3.0 Flash: 124B total parameters, 5.1B active per token (MoE architecture)
256K token context window
Free access through August 3rd via AI Gateway
Designed for agentic inference, coding agents, document work, long-context interactions
Available via Vercel AI SDK with model identifier: inclusionai/ling-3.0-flash-free
AI Gateway pricing: no markup, no platform fee on inference, includes BYOK support
Thinking and non-thinking modes supported
Ling 3.0 Flash: 124B total parameters, ~5.1B active per token
Mixture-of-Experts architecture optimized for token efficiency
Free access through August 3rd, 2026
Available via Vercel AI Gateway with no platform markup or inference fees
Targeting high-frequency agentic workflows, coding agents, document work, long-context multi-turn interactions
Supports thinking and non-thinking modes
AI Gateway includes custom reporting, Zero Data Retention, budgets for API keys, routing rules
Go to the source
Vercel Blogvercel.com
Publisher excerpt: from Ant Group is now available on AI Gateway.Ling 3.0 Flash The model is free to use for the next three weeks, through August 3rd. Ling 3.0 Flash is a Mixture-of-Experts model with 124B total parameters and about 5.1B active per token. It has a 256K token context window and runs in thinking and…