Opening film
Similar motions. Different forces.
Gentle and firm actions can look almost identical, yet apply very different forces. Opt2VLA lets people specify both what the robot should do and how strongly it should interact.
Illustrated simulation rollouts. 1080p preview, 91 seconds.
Rendering and data notes
Successful contact-rich manipulation requires more than moving a hand to the right place. Too little force can let an object slip or fail to move it; too much can damage a delicate object or destabilize the robot. The appropriate force depends on the task, object properties, and safety constraints, even when the visible motion is similar. A humanoid must regulate these interactions while maintaining whole-body balance.
Opt2VLA makes desired contact force an explicit command alongside motion, rather than leaving it to emerge from motion tracking alone. A single VLA policy predicts motion and contact-force commands from language, vision, robot state, and measured joint-torque feedback. Task-specific force-aware whole-body controllers execute those commands, connecting instructions such as “gently” and “strongly” to how the robot physically interacts with its surroundings.
The film illustrates this idea using recorded closed-loop simulation rollouts. Blue denotes issued force; amber denotes measured force. Robot motion and force values come from the recordings. Arrows share a length scale, with schematic outward directions representing scalar force magnitudes.
Materials, cloth, stains, and scene dressing illustrate why different force levels matter; they were added for the film and were not observed by the policy. Physical loads are unchanged. These scenes do not test material or weight inference, or cleaning effectiveness.