By Interestana AI Editorial — AI-drafted, human-overseen. How we report
NVIDIA Researchers Develop Physis-Lang for Physics-Accurate Video
NVIDIA researchers, in collaboration with MIT and the University of Oxford, have introduced Physis-Lang, an open, self-evolving framework designed to enhance the physical accuracy of video world models. This novel approach posits that physical reasoning can be integrated through language itself, rather than relying on additional visual, latent, or numerical signals. Physis-Lang treats physical language as a shared, optimizable representation that simultaneously drives data curation, model training, and inference processes.
The core innovation of Physis-Lang lies in its ability to generate captions that not only describe the visual events in a video but also explain the underlying physical principles governing those events. Conventional video captions typically focus on observable actions, such as 'butter melts as the temperature rises,' without detailing the physics of heat transfer or gravity. Physis-Lang augments these base captions with a dedicated `physics_reasoning` field. This field explicitly outlines entities involved, causal relationships, interactions, governing physical laws, temporal progression, and the resulting effects within a scene. Furthermore, the framework generates a scene-specific `physics_negative_prompt` that details plausible but physically impossible outcomes, such as a stone floating on water. This negative prompt is then utilized as negative conditioning during the inference stage to steer the model away from generating physically inaccurate content.
The self-evolving caption loop is central to Physis-Lang's development. In this process, the captioning model itself remains frozen, while its instructions are iteratively refined. A GPT-5.5 captioner is employed to generate captions for a fixed development set comprising 20 videos and 273 human-verified physical assertions. A physics-aware critic, utilizing Gemini-3.1-Pro, evaluates these captions. The critic assesses two primary dimensions: Precision, which involves breaking down captions into atomic claims and verifying each claim against established physics, and a second dimension that identifies claim-level failures. An evolution agent then analyzes the scores and identified failures to rewrite the prompt for the GPT-5.5 captioner, thereby continuously improving the quality and accuracy of the generated physics explanations.
On the public Physics-IQ Verified leaderboard, as of a snapshot taken on September 29, 2026, Physis-Lang demonstrated significant performance gains. The Cosmos3-Super version of the framework achieved a top-ranking score of 48.2 ± 1.4. The Cosmos3-Nano version secured the second position with a score of 43.3 ± 1.5. These results indicate that Physis-Lang, by integrating physics reasoning directly into the language used to describe video content, significantly surpasses existing video world models in physical accuracy, outperforming models like Veo 3.1 on physics benchmarks. The framework's open nature and self-evolving mechanism suggest a path towards more physically coherent and believable AI-generated video content.
Original source — read the full reporting at the publisher:
Read on MarkTechPostGet the weekly AI digest
AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.