FrontierThe story, in brief

NVIDIA Researchers Introduce Physis-Lang: Self-Evolving Physical Language That Lifts Cosmos 3 Past Veo 3.1 on Physics Benchmarks

NVIDIA's Physis-Lang beats Google Veo 3.1 on physics fidelity. The fix: language itself, not more visual data.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

NVIDIA researchers propose that physical language—a learned, shared representation—can ground video world models in correct physics better than latent or numerical signals alone. Cosmos 3 outperforms Veo 3.1 on physics benchmarks. Practitioners building video generation or physics-constrained synthesis should track whether this approach generalizes beyond the benchmark.

The key facts

6 to know
  1. Physis-Lang framework treats physical language as optimizable shared representation

  2. Cosmos 3 lifts above Veo 3.1 on physics benchmarks (specific benchmark names and scores not disclosed in excerpt)

  3. Collaboration: NVIDIA, MIT, University of Oxford

  4. Problem addressed: video world models rendering non-physical artifacts (butter spreading like paint, balls passing through walls)

  5. Mechanism: language-based grounding rather than visual, latent, or numerical signals

  6. Source: MarkTechPost (secondary reporting; original research paper or NVIDIA announcement not linked in excerpt)

Go to the source

MarkTechPostmarktechpost.com

Publisher excerpt: Video world models can render convincing clips that still break physics. Butter spreads like paint. Balls pass through walls. A team from NVIDIA, MIT and the University of Oxford argues the fix can come from language itself, not from extra visual, latent or numerical signals. Their framework,…
Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier