ChipsThe story, in brief

Incredibly Fast BLOOM Inference with DeepSpeed and Accelerate

Open-source BLOOM model now runs 5-10x faster. Here's how Hugging Face and Microsoft did it.

Paper-cut illustration of an amber microchip with circuit paths extending into a row of data-center cabinets.
The infrastructure powering AI.AI illustration by KeyNews
The KeyNews take

Why it matters

Inference optimization tools (DeepSpeed + Accelerate) dramatically lower the operational cost and latency barrier for running large language models, making open-source model deployment viable at scale for more organizations.

The key facts

10 to know
  1. BLOOM model inference optimization via DeepSpeed and Accelerate

  2. Focus on inference speed improvements for open-source LLM deployment

  3. Joint work between Hugging Face and Microsoft

  4. Published September 2022 — foundational period for open-source LLM infrastructure

  5. Reduces compute barriers for model serving and deployment

  6. BLOOM inference optimization via DeepSpeed and Accelerate

  7. Focus on inference speed and cost efficiency

  8. Collaboration between Hugging Face and Microsoft

  9. Published September 2022 (historical but foundational)

  10. Open-source inference optimization as infrastructure play

Go to the source

Hugging Face Bloghuggingface.co

Read original report
Back to today's editionMore chips news

The wider picture

View all
Paper-cut illustration of an amber microchip with circuit paths extending into a row of data-center cabinets.
AI illustration by KeyNews
Chips01

Why confidential computing is essential for enterprise AI

As enterprises push AI into sensitive domains (healthcare, finance, government), protecting data during processing — not just at rest or in transit — is shifting from a nice-to-have security feature to a prerequisite for deployment. HPE and NVIDIA are positioning confidential computing as foundational infrastructure for sovereign AI.

CIO
Paper-cut illustration of an amber microchip with circuit paths extending into a row of data-center cabinets.
AI illustration by KeyNews
Chips02

Beyond the limits of air: Why liquid cooling is becoming a strategic imperative for AI

Liquid cooling is shifting from niche to strategic necessity as GPU power density explodes. For practitioners building or deploying AI at scale, this isn't optional—it's a 18–24 month ROI play that unlocks denser racks, lower OpEx, and competitive advantage. The buildout architecture is changing.

CIO
Paper-cut illustration of an amber microchip with circuit paths extending into a row of data-center cabinets.
AI illustration by KeyNews
Chips03

Singapore’s Nexstrom wants to bring 2D semiconductors to chip fabs

2D materials promise better performance-per-watt for next-gen AI accelerators. Nexstrom's equipment funding signals the transition from lab to foundry—a potential shift in the compute buildout.

TechCrunch Startups