Yash V. Prabhu
MIT MGAIC | MIT Generative AI Impact Research and Innovation Scholar
Bridging the Humanoid Manipulation Sim2Real Gap via Sim-Real Co-Training
2026–2027
Electrical Engineering and Computer Science
- Robotics
Pulkit Agrawal
Dexterous humanoid manipulation remains a central problem in robotics because it requires delicate contact-rich interaction with complex real-world environments. While reinforcement learning (RL) has enabled major progress in dynamic humanoid locomotion through large-scale simulation, dexterous manipulation presents a significant exploration challenge: policies must discover long-horizon contact behaviors from sparse rewards, while designing dense reward functions for complex tasks can be difficult and brittle. Even when effective behaviors are learned in simulation, transfer to physical hardware remains difficult due to the sim-to-real gap in contact dynamics. Conversely, imitation learning (IL) from real-world demonstrations provides strong behavioral priors and improved transfer, but is limited by the cost, quality, and diversity of human-collected data. This project proposes a sim-real co-training framework that combines RL in simulation with IL from real-world whole-body teleoperation demonstrations. By jointly optimizing RL objectives and behavior-cloning losses, the approach aims to improve exploration efficiency, enhance sim-to-real transfer, and enable policies to exceed the capabilities of their demonstrators. Training will use massively parallel simulation environments together with teleoperation data collected on a Unitree G1 humanoid robot. The framework will be evaluated on dexterous humanoid manipulation tasks in simulation and the real world, with performance measured by task success, robustness, generalization, and transfer quality. The goal is to learn policies that explore efficiently, transfer reliably, exceed their demonstrations, and generalize across dexterous humanoid manipulation tasks.
I believe that general-purpose robots have the potential to significantly reduce manual burdens placed on humans, allowing people to devote more of their time to what they find meaningful. As advances in embodied intelligence continue to accelerate, I am excited to contribute to the development of intelligent robotic systems through SuperUROP and to further develop as a researcher in the process.
