FrontierAugust 2, 2026via MarkTechPost
Thinking Machines Lab Releases Inkling-Small: A 276B Total, 12B Active Open Weights Multimodal MoE Model
Why it matters
Open-weight MoE efficiency milestone: a multimodal model achieving parity with a larger peer via sparsity and quantization, lowering the bar for practitioners to run capable models on consumer/mid-tier hardware.
Key signals
- Inkling-Small: 276B total parameters, 12B active (MoE sparse)
- Multimodal (vision + language) capability
- Matches Inkling (full-size version) performance at 1/4 scale
- NVFP4 quantization checkpoint
- Runs on single NVIDIA B300 GPU
- Open weights release
- Published: August 2, 2026
The hook
276B parameters, 12B active—Thinking Machines Lab's Inkling-Small matches its flagship at a quarter the size and fits on a single B300.
Inkling-Small matches Inkling at a quarter the size, and its NVFP4 checkpoint runs on one NVIDIA B300 GPU