NVIDIA Researchers Introduce Physis-Lang: Self-Evolving Physical Language That Lifts Cosmos 3 Past Veo 3.1 on Physics Benchmarks
NVIDIA's Physis-Lang beats Google Veo 3.1 on physics fidelity. The fix: language itself, not more visual data.

Why it matters
NVIDIA researchers propose that physical language—a learned, shared representation—can ground video world models in correct physics better than latent or numerical signals alone. Cosmos 3 outperforms Veo 3.1 on physics benchmarks. Practitioners building video generation or physics-constrained synthesis should track whether this approach generalizes beyond the benchmark.
The key facts
6 to knowPhysis-Lang framework treats physical language as optimizable shared representation
Cosmos 3 lifts above Veo 3.1 on physics benchmarks (specific benchmark names and scores not disclosed in excerpt)
Collaboration: NVIDIA, MIT, University of Oxford
Problem addressed: video world models rendering non-physical artifacts (butter spreading like paint, balls passing through walls)
Mechanism: language-based grounding rather than visual, latent, or numerical signals
Source: MarkTechPost (secondary reporting; original research paper or NVIDIA announcement not linked in excerpt)
Go to the source
MarkTechPostmarktechpost.com
Publisher excerpt: Video world models can render convincing clips that still break physics. Butter spreads like paint. Balls pass through walls. A team from NVIDIA, MIT and the University of Oxford argues the fix can come from language itself, not from extra visual, latent or numerical signals. Their framework,…