ToolsThe story, in brief

Show HN: Smart model routing directly in Claude, Codex and Cursor

40% cost savings. Weave just shipped a model router that intelligently picks the right AI model for each coding task—and it's available now.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

As frontier model costs climb (Opus 4.7's tokenizer changes drove Weave's expenses up), intelligent model routing is becoming table-stakes for AI-first teams. This is a working solution shipping with real cost/quality data.

The key facts

6 to know
  1. 40% token cost reduction with no quality loss

  2. Routes requests across Opus 4.8, GPT 5.5, DeepSeek V4, GLM 5.2, Kimi K2.6

  3. Trained RL routing model on tens of thousands of agent traces

  4. Integrates with Claude Code, Codex, Cursor as drop-in endpoint

  5. Source-available under Elastic License 2.0; hosted version at weaverouter.com

  6. Built after Opus 4.7 tokenizer changes spiked internal AI coding costs

Go to the source

Hacker Newsgithub.com

Publisher excerpt: We built a model router that plugs into coding agents (e.g. Claude Code, Codex, Cursor, etc.) and intelligently sends requests to the best model to serve them. Here's a quick demo of running it locally: At Weave, we write ~all our code with AI, and it's been getting more expensive. This came to a…
Read original report
Back to today's editionMore tools news

Keep reading

Related stories

More from Tools