From Where Things Are to What They’re For: Benchmarking Spatial–Functional Intelligence for Multimodal LLMs
Apple just raised the bar on spatial reasoning. New benchmark shows multimodal LLMs are stuck at 'where' — not 'why.'

Why it matters
Apple's new SFI-Bench exposes a critical gap in multimodal AI: existing models excel at geometric perception but fail at functional reasoning — a capability essential for embodied agents. This benchmark becomes the new standard for evaluating spatial intelligence in production systems.
The key facts
7 to knowSFI-Bench: 1,700+ video-based benchmark questions
Data sourced from diverse egocentric indoor video scans
Benchmark differentiates geometric perception from functional understanding
Designed to probe higher-order cognitive abilities for grounded intelligence
Addresses limitation of prior benchmarks like VSI-Bench
Published by Apple ML Research (May 2026)
Focus on multimodal LLM evaluation for spatial reasoning
Go to the source
Apple Machine Learningmachinelearning.apple.com
Publisher excerpt: True spatial intelligence for multimodal agents transcends low-level geometric perception, evolving from knowing where things are to understanding what they are for. While existing benchmarks, such as VSI-Bench, effectively evaluate this foundational geometric stage, they fall short of probing the…