Maryna Bohdan
MIT MGAIC | MIT Generative AI Impact Research and Innovation Scholar
Language Models as Physics Compilers for Process Rewards in Generative Video Models
2026–2027
Physics; Electrical Engineering and Computer Science
- AI and Machine Learning
- Physics
William T. Freeman
As generative video models become increasingly realistic, physical consistency remains one of the key challenges limiting their reliability. For example, a video may appear convincing while still breaking the physical rules implied by the scene, such as gravity, momentum, collisions, or object permanence. This project investigates how language models can be used as “physics compilers” for video generation. Given a prompt or scene description, a language model would infer the relevant physical assumptions and convert them into equations or simulator-based reward functions. These rewards would provide process-level feedback by measuring whether frame-to-frame motion is physically plausible. The goal of this work is to connect symbolic physical reasoning with neural video generation and contribute to models that are not only visually realistic, but also physically grounded.
I am participating in SuperUROP because I am interested in the gap between videos that look realistic and videos that are physically plausible. As generative video models become more capable, I want to understand whether they are learning meaningful structure about the world or mainly producing convincing visual patterns. Through this project, I hope to grow as a researcher while contributing to work that making generative video models more physically grounded.
