Xiaomi MiMo and TileRT Push a 1-Trillion-Parameter Model Past 1000 Tokens Per Second on Commodity GPUs
1,000 tokens per second. On a single 8-GPU node. Xiaomi just proved trillion-parameter models don't need custom silicon.

Why it matters
Xiaomi's TileRT serving optimization fundamentally shifts the economics of LLM deployment. If commodity GPUs can now handle trillion-parameter inference at scale, the infrastructure moat for custom-chip builders narrows significantly.
The key facts
5 to knowMiMo-V2.5-Pro-UltraSpeed achieves 1000+ tokens/second throughput
1-trillion-parameter model running on single 8-GPU commodity node
Inference optimization via TileRT serving mode
Published June 2026 (future-dated; UNVERIFIED)
Deployment target: commodity GPUs (not custom silicon)
Go to the source
MarkTechPostmarktechpost.com
Publisher excerpt: Xiaomi's MiMo team, with TileRT, released MiMo-V2.5-Pro-UltraSpeed, a serving mode for the MiMo-V2.5-Pro model. It decodes over 1000 tokens per second on a 1-trillion-parameter model using a single 8-GPU commodity node.