FrontierThe story, in brief

Meet Needle 2: An Open 45M-Parameter Tool-Calling Model That Ships as a 14MB Binary and Runs a Full Session in 28MB of RAM

45M parameters. 14MB binary. Needle 2 runs a full AI session on hardware with zero GPUs—and leads tool-calling benchmarks doing it.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

A new open-weight model challenges the assumption that capable AI requires scale and silicon. Needle 2's efficiency (sub-30MB RAM, no accelerators) matters for practitioners building on-device, edge, and resource-constrained deployments — and signals a shift in how frontier labs think about capability-per-watt.

The key facts

8 to know
  1. Needle 2: 45M parameters

  2. Binary size: 14MB

  3. Session RAM footprint: ~28MB

  4. No GPU or NPU required

  5. Leads Seal-Tools benchmark splits

  6. Open-weight release

  7. Vendor: Cactus Compute

  8. Use cases: tool calling, device use, structured extraction

Go to the source

MarkTechPostmarktechpost.com

Publisher excerpt: Cactus Compute released Needle 2, an open 45M-parameter model for tool calling, device use, and structured extraction. The full model is a single 14MB binary that runs a session in about 28MB of RAM. It leads both Seal-Tools splits while targeting hardware with no GPU and no NPU. The post Meet…
Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier