FrontierAugust 2, 2026via MarkTechPost

Thinking Machines Lab Releases Inkling-Small: A 276B Total, 12B Active Open Weights Multimodal MoE Model

Why it matters

Open-weight MoE efficiency milestone: a multimodal model achieving parity with a larger peer via sparsity and quantization, lowering the bar for practitioners to run capable models on consumer/mid-tier hardware.

Key signals

  • Inkling-Small: 276B total parameters, 12B active (MoE sparse)
  • Multimodal (vision + language) capability
  • Matches Inkling (full-size version) performance at 1/4 scale
  • NVFP4 quantization checkpoint
  • Runs on single NVIDIA B300 GPU
  • Open weights release
  • Published: August 2, 2026

The hook

276B parameters, 12B active—Thinking Machines Lab's Inkling-Small matches its flagship at a quarter the size and fits on a single B300.

Inkling-Small matches Inkling at a quarter the size, and its NVFP4 checkpoint runs on one NVIDIA B300 GPU

The week's key stories, every Friday.

For practitioners and enthusiasts — free, in your inbox.

Free forever. No spam.