AI agents build 3D scenes from photos but have no idea if they got it right
GPT-6 Astra built 3D scenes from photos with 53% accuracy. Problem: agents can't tell if they got it right—they fail at self-evaluation worse than a coin flip.

Why it matters
LEGO-Anything demonstrates agent capability at a concrete task (3D reconstruction), but exposes a critical reliability gap: agents lack intrinsic ability to validate their own geometric accuracy. For production deployments, this means external verification overhead.
The key facts
5 to knowLEGO-Anything: single-photo to editable Blender code for 3D scenes
GPT-6 Astra leads benchmark with up to 53% reconstruction accuracy
Critical weakness: all tested agents cannot judge geometric accuracy better than coin flip (50% baseline)
Benchmark introduced; no independent validation cited
Unpublished paper or dataset; limited methodological detail in article
Go to the source
The Decoderthe-decoder.com
Publisher excerpt: A new approach called LEGO-Anything turns single photos into editable Blender code for 3D scenes. GPT-6 Astra leads the accompanying benchmark with up to 53 percent reconstruction accuracy. The biggest weakness across all tested agents is that they can't judge their own geometric accuracy any…