Yibo Cheng
MIT EECS | Boeing Undergraduate Research and Innovation Scholar
Sensor Fusion for Contact Mode Inference
2026–2027
Electrical Engineering and Computer Science
- Robotics
Nicholas Roy
Paul Liang
Sequential robotic manipulation depends on correctly tracking contact mode: each action makes or breaks contact, and a single early failure-such as missing a grasp on a drawer handle-invalidates every downstream step. Yet inferring the current contact configuration from any single sensor is unreliable. Force-torque readings are ambiguous about which surfaces are actually in contact, and vision from a single viewpoint is frequently occluded by the arm or the object itself. This project develops a method to jointly infer object pose and contact mode by fusing visual and force-torque sensing within a particle filter, where each particle maintains its own Drake simulation to predict the expected contact response. We combine three likelihood factors into a single scalar energy E(ξ): a vision factor penalizing point-cloud disagreement between observed depth and the object mesh under the hypothesized pose, a force-torque factor penalizing mismatch between measured and predicted wrenches at the estimated contact configuration, and a penetration factor penalizing geometric interpenetration between object and hand. Using exp(−E(ξ)) as the observation likelihood yields a posterior over object pose that is simultaneously consistent with visual observations, contact measurements, and physical plausibility. Building on an existing force-torque particle filter and a Grounded-SAM-2 + FoundationPose vision pipeline, we aim to evaluate the fused estimator on a real robot and assess whether the resulting contact and pose estimates are accurate enough to support reliable manipulation planning.
I’m drawn to the problem of giving robots a reliable sense of their own contact with the world, because it sits right at the intersection of physical reasoning, perception, and planning that first pulled me toward robotics. What excites me about this project is that neither vision nor touch alone is enough-the interesting work is in figuring out how to make ambiguous, noisy, and heterogeneous signals agree with each other and with physics. SuperUROP gives me the chance to take a problem from a clean simulation result all the way to a real robot, where the messiness of actual sensor data forces the ideas to be honest. I want to spend a sustained stretch of time owning a research question end to end, learning to move between the theory of Bayesian estimation and the engineering reality of getting a system to work on hardware
