ToolsJuly 23, 2026via Vercel Blog

Ling 3.0 Flash is now available on AI Gateway

Why it matters

Ant Group's latest open model is now accessible via Vercel's unified API layer, lowering friction for developers building multi-step agent applications. The free tier and token-efficient design (5.1B active params) signal competitive pressure on inference costs in the agentic inference market.

Key signals

  • Ling 3.0 Flash: 124B total parameters, 5.1B active per token (MoE architecture)
  • 256K token context window
  • Free access through August 3rd via AI Gateway
  • Designed for agentic inference, coding agents, document work, long-context interactions
  • Available via Vercel AI SDK with model identifier: inclusionai/ling-3.0-flash-free
  • AI Gateway pricing: no markup, no platform fee on inference, includes BYOK support
  • Thinking and non-thinking modes supported
  • Ling 3.0 Flash: 124B total parameters, ~5.1B active per token
  • Mixture-of-Experts architecture optimized for token efficiency
  • Free access through August 3rd, 2026
  • Available via Vercel AI Gateway with no platform markup or inference fees
  • Targeting high-frequency agentic workflows, coding agents, document work, long-context multi-turn interactions
  • Supports thinking and non-thinking modes
  • AI Gateway includes custom reporting, Zero Data Retention, budgets for API keys, routing rules

The hook

Ant Group's Ling 3.0 Flash hits Vercel's AI Gateway. Free for 3 weeks. 124B MoE built for agentic workflows at production scale.

from Ant Group is now available on AI Gateway.Ling 3.0 Flash The model is free to use for the next three weeks, through August 3rd. Ling 3.0 Flash is a Mixture-of-Experts model with 124B total parameters and about 5.1B active per token. It has a 256K token context window and runs in thinking and non-thinking modes. Ling 3.0 Flash is built for token-efficient agentic inference at production scale, doing more work within tighter token, latency, and cost budgets across multi-step agent runs. The model targets high-frequency agentic workflows, coding agents, document work, and long-context multi-turn interactions. To use Ling 3.0 Flash, set model to in the :inclusionai/ling-3.0-flash-freeAI SDK AI Gateway provides a unified API for calling models, tracking usage and cost, and configuring retries, failover, and performance optimizations for higher-than-provider uptime. It includes built-in , , , , and more.custom reportingZero Data Retention supportbudgets for API keysrouting rules AI Gateway reflects provider pricing with no markup and does not charge a platform fee on inference, including on (BYOK) requests.Bring Your Own Key Try Ling 3.0 Flash in the .model playground

The week's key stories, every Friday.

For practitioners and enthusiasts — free, in your inbox.

Free forever. No spam.