ChipsThe story, in brief

Pinterest Wants Smarter Visual Search Without the Huge AI Computing Bill

Pinterest cuts its visual search inference bill by optimizing on Nvidia Blackwell — a blueprint for how platforms are squeezing AI compute costs.

Paper-cut illustration of an amber microchip with circuit paths extending into a row of data-center cabinets.
The infrastructure powering AI.AI illustration by KeyNews
The KeyNews take

Why it matters

A major consumer platform is deploying inference optimization (Dynamo) on new silicon to reduce latency and cost of multimodal search at scale — a real-world case study in the compute economics reshaping cloud AI workloads.

The key facts

10 to know
  1. Pinterest using Nvidia Blackwell GPUs for multimodal AI search

  2. Dynamo optimization framework deployed to reduce inference latency

  3. Focus on reducing AI computing costs while maintaining visual search capability

  4. Assistant feature receiving enhanced visual context processing

  5. Pinterest using Nvidia Blackwell GPUs for visual search

  6. Dynamo inference acceleration deployed

  7. Focus on reducing inference latency and cost

  8. Multimodal AI search optimization

  9. Production deployment (not pilot)

  10. Visual context enhancement for Pinterest Assistant

Go to the source

TechRepublictechrepublic.com

Publisher excerpt: Pinterest is using Nvidia Blackwell GPUs and Dynamo to speed up multimodal AI search, reduce inference latency, and give its Assistant more visual context.
Read original report
Back to today's editionMore chips news

The wider picture

View all
Paper-cut illustration of an amber microchip with circuit paths extending into a row of data-center cabinets.
AI illustration by KeyNews
Chips01

NVIDIA Launches DSX Ready to Qualify Power and Cooling Products for AI Factories

As AI compute density increases, power and cooling constraints are reshaping data-center buildout decisions. NVIDIA's qualification framework helps operators match infrastructure to workload architecture — a critical lever in the AI factory economics game.

NVIDIA Blog
Paper-cut illustration of an amber microchip with circuit paths extending into a row of data-center cabinets.
AI illustration by KeyNews
Chips02

How BMW Group detects cost anomalies across 14,000 cloud accounts

BMW's serverless anomaly-detection pipeline illustrates how enterprises are automating cloud-cost governance at scale—a real use case for ML in the compute buildout that practitioners managing multi-cloud infrastructure should know about.

AWS Machine Learning Blog
Paper-cut illustration of an amber microchip with circuit paths extending into a row of data-center cabinets.
AI illustration by KeyNews
Chips03

US data centres ‘are short six NYCs of electricity’

AI's explosive compute demands have outpaced power grid capacity. Data-centre operators and their investors now face a physics problem, not just an economics one — and it will reshape where AI gets built.

Financial Times Technology