ChipsSeptember 10, 2026via MarkTechPost

NVIDIA Details BioNeMo Inference Runtime (BioIR): 2.90x Higher Boltz-2 Folding Throughput and 58.5K Residues per GPU-Hour on 8xH100

Why it matters

NVIDIA is optimizing GPU utilization for a critical AI workload (biomolecular simulation) with specialized inference software. This matters to practitioners deploying compute-intensive scientific AI and signals how chip makers compete beyond raw FLOPS—through runtime software and domain-specific optimization.

Key signals

  • BioNeMo Inference Runtime (BioIR) achieves 2.90x throughput gain on Boltz-2 protein folding
  • 58.5K residues per GPU-hour on 8xH100 vs. 20.2K with torch-compiled baseline
  • Benchmark: 1,000 human dimer targets
  • Three-layer optimization: custom kernels, CUDA Graph capture, Ray-based replica scaling
  • Already deployed: ~31 million protein complex candidates generated across 4,777 proteomes for AlphaFold Database expansion
  • Stays in plain PyTorch ecosystem

The hook

2.90x. That's how much faster NVIDIA's new inference runtime makes protein folding on H100s—and it's already scaled to 31 million structures.

NVIDIA has detailed BioNeMo Inference Runtime (BioIR), a Python library that accelerates biomolecular structure-prediction models on NVIDIA GPUs while staying in plain PyTorch. In a matched benchmark on 1,000 human dimer targets across 8xH100 GPUs, BioIR-accelerated Boltz-2 delivered 58.5K successfu

The week's key stories, every Friday.

For practitioners and enthusiasts — free, in your inbox.

Free forever. No spam.