AMD acquires Taalas, a startup that bakes AI models directly into silicon
AMD just bought the company that bakes model weights into silicon—16,000 tokens/sec per user, but locked to one model per chip.

Why it matters
AMD is making a strategic bet on model-in-silicon inference as a competitive edge against Nvidia, signaling a shift toward purpose-built hardware that trades flexibility for extreme speed. This is how the compute buildout evolves when you own both the chip and the AI stack.
The key facts
6 to knowAMD acquires Taalas (Canadian startup)
Taalas hard-codes model weights directly into inference chips
Demo chip achieved 16,000 tokens/second per user running Llama 3.1-8B
Trade-off: extreme speed but single-model lock-in per chip
Google reportedly pursuing similar approach for Gemini
Published: August 7, 2026
Go to the source
The Decoderthe-decoder.com
Publisher excerpt: AMD is buying Canadian startup Taalas, which hard-codes model weights directly into inference chips. That makes them extremely fast but locks each chip to a single model. A demo chip hit over 16,000 tokens per second per user running Llama 3.1-8B. Google is reportedly working on a similar approach…