ChipsThe story, in brief

14× faster embeddings: how we rebuilt the ONNX path in Manticore

14× faster. That's what Manticore just unlocked by rebuilding their ONNX inference path.

Paper-cut illustration of an amber microchip with circuit paths extending into a row of data-center cabinets.
The infrastructure powering AI.AI illustration by KeyNews
The KeyNews take

Why it matters

Infrastructure optimization in vector search and embedding inference directly impacts AI application latency and operational costs. A 14× speedup in ONNX execution is the kind of incremental but production-critical engineering that compounds across thousands of deployed AI systems.

The key facts

9 to know
  1. 14× performance improvement in ONNX embedding inference

  2. Manticore rebuilt ONNX execution path for optimization

  3. Focus on vector search infrastructure efficiency

  4. Published Jul 03 2026 on Manticore's official technical blog

  5. 14× performance improvement in embedding inference

  6. ONNX runtime optimization

  7. Focus on vector search infrastructure

  8. Published July 2026

  9. Low engagement (3 points, 0 comments on HN) suggests niche technical audience

Go to the source

Hacker Newsmanticoresearch.com

Publisher excerpt: Article URL: Comments URL: Points: 3 # Comments: 0
Read original report
Back to today's editionMore chips news

The wider picture

View all
Paper-cut illustration of an amber microchip with circuit paths extending into a row of data-center cabinets.
AI illustration by KeyNews
Chips01

Google Adds Cycle-Level Kernel Profiling to XProf

A developer-facing tooling improvement that directly enables better TPU utilization and kernel optimization. Practitioners building custom Pallas kernels can now see exactly where cycles are spent, shifting from guesswork to data-driven tuning.

InfoQ AI/ML
Paper-cut illustration of an amber microchip with circuit paths extending into a row of data-center cabinets.
AI illustration by KeyNews
Chips02

Civo unveils first of 40 planned edge data center sites across UK

Edge compute infrastructure is becoming critical for low-latency AI inference and agentic workloads. Civo's distributed network strategy reflects growing demand for regional AI compute capacity outside centralized cloud zones — a structural shift in how AI workloads are deployed.

ITPro
Paper-cut illustration of an amber microchip with circuit paths extending into a row of data-center cabinets.
AI illustration by KeyNews
Chips03

China reviews dependence on Broadcom switches in data centres

China is auditing its reliance on foreign networking hardware for AI data centers as part of a broader push to build domestic alternatives. This reshapes global compute buildout economics and chip supply chains at a moment when AI capacity is the competitive moat.

Financial Times Technology