ChipsThe story, in brief

Xiaomi MiMo and TileRT Push a 1-Trillion-Parameter Model Past 1000 Tokens Per Second on Commodity GPUs

1,000 tokens per second. On a single 8-GPU node. Xiaomi just proved trillion-parameter models don't need custom silicon.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

Xiaomi's TileRT serving optimization fundamentally shifts the economics of LLM deployment. If commodity GPUs can now handle trillion-parameter inference at scale, the infrastructure moat for custom-chip builders narrows significantly.

The key facts

5 to know
  1. MiMo-V2.5-Pro-UltraSpeed achieves 1000+ tokens/second throughput

  2. 1-trillion-parameter model running on single 8-GPU commodity node

  3. Inference optimization via TileRT serving mode

  4. Published June 2026 (future-dated; UNVERIFIED)

  5. Deployment target: commodity GPUs (not custom silicon)

Go to the source

MarkTechPostmarktechpost.com

Publisher excerpt: Xiaomi's MiMo team, with TileRT, released MiMo-V2.5-Pro-UltraSpeed, a serving mode for the MiMo-V2.5-Pro model. It decodes over 1000 tokens per second on a 1-trillion-parameter model using a single 8-GPU commodity node.
Read original report
Back to today's editionMore chips news

Keep reading

Related stories

More from Chips