Inkling Small from Thinking Machines is now available on AI Gateway
Inkling Small hits AI Gateway: quarter the size, same reasoning chops. Coding agents just got cheaper.

Why it matters
A smaller, more efficient reasoning model is now available through Vercel's AI Gateway, lowering the cost barrier for practitioners building agentic workflows and coding tools. The native multimodal reasoning and controllable thinking-effort feature address real developer friction around inference costs.
The key facts
14 to knowInkling Small achieves performance comparable to larger Inkling model at ~25% of the size
Native reasoning over audio and images
Controllable thinking effort: trade quality against cost and latency
Compatible with Zero Data Retention (ZDR) for privacy-sensitive deployments
AI Gateway routes to providers that delete prompts/responses per request
No platform fee on inference, including BYOK requests
Optimized for coding agents and tool-use workflows
Available on AI Gateway playground for testing
Inkling Small reaches performance of larger Inkling at ~1/4 the size
Controllable thinking effort (quality/cost/latency tradeoff)
Compatible with Zero Data Retention for privacy-sensitive workflows
Available in Vercel AI Gateway with no platform markup on inference
Supports coding agents and tool-use workflows
Visual capabilities: programmatic cropping, zooming, image inspection for documents/charts
Go to the source
Vercel Blogvercel.com
Publisher excerpt: from Thinking Machines is now available on AI Gateway.Inkling Small Inkling Small reaches performance comparable to the larger Inkling model at about a quarter of the size, using much less compute per task. It is a broad generalist with native reasoning over audio and images, and it holds up well…