MoneyThe story, in brief

AMD acquires Taalas, a startup that bakes AI models directly into silicon

AMD just bought the company that bakes model weights into silicon—16,000 tokens/sec per user, but locked to one model per chip.

Paper-cut illustration of amber paths carrying capital toward a small coral research venture between larger buildings.
Capital and the next generation of AI ventures.AI illustration by KeyNews
The KeyNews take

Why it matters

AMD is making a strategic bet on model-in-silicon inference as a competitive edge against Nvidia, signaling a shift toward purpose-built hardware that trades flexibility for extreme speed. This is how the compute buildout evolves when you own both the chip and the AI stack.

The key facts

6 to know
  1. AMD acquires Taalas (Canadian startup)

  2. Taalas hard-codes model weights directly into inference chips

  3. Demo chip achieved 16,000 tokens/second per user running Llama 3.1-8B

  4. Trade-off: extreme speed but single-model lock-in per chip

  5. Google reportedly pursuing similar approach for Gemini

  6. Published: August 7, 2026

Go to the source

The Decoderthe-decoder.com

Publisher excerpt: AMD is buying Canadian startup Taalas, which hard-codes model weights directly into inference chips. That makes them extremely fast but locks each chip to a single model. A demo chip hit over 16,000 tokens per second per user running Llama 3.1-8B. Google is reportedly working on a similar approach…
Read original report
Back to today's editionMore money news

Keep reading

Related stories

More from Money