MoneyAugust 7, 2026via The Decoder

AMD acquires Taalas, a startup that bakes AI models directly into silicon

Why it matters

AMD is making a strategic bet on model-in-silicon inference as a competitive edge against Nvidia, signaling a shift toward purpose-built hardware that trades flexibility for extreme speed. This is how the compute buildout evolves when you own both the chip and the AI stack.

Key signals

  • AMD acquires Taalas (Canadian startup)
  • Taalas hard-codes model weights directly into inference chips
  • Demo chip achieved 16,000 tokens/second per user running Llama 3.1-8B
  • Trade-off: extreme speed but single-model lock-in per chip
  • Google reportedly pursuing similar approach for Gemini
  • Published: August 7, 2026

The hook

AMD just bought the company that bakes model weights into silicon—16,000 tokens/sec per user, but locked to one model per chip.

AMD is buying Canadian startup Taalas, which hard-codes model weights directly into inference chips. That makes them extremely fast but locks each chip to a single model. A demo chip hit over 16,000 tokens per second per user running Llama 3.1-8B. Google is reportedly working on a similar approach f

The week's key stories, every Friday.

For practitioners and enthusiasts — free, in your inbox.

Free forever. No spam.