What if robot hands could feel?
Robot vision is advancing quickly. We ask what happens when robots can also feel contact directly, and use it to choose their next move.
INTERACT ROBOTICS · Insights · Reviewed September 7, 2026
When seeing isn’t feeling.
Imagine fitting a small connector through a thick glove. It looks almost aligned, yet the direction of a catch and the quality of the fit are things your fingertips would normally tell you. With that sense dulled, familiar force adjustments take more checking.
“Anesthetized hands” is our metaphor for this missing information. A robot without direct force and tactile observations must infer what happened after contact from vision, position, and experience. It does not equate robots with people who have lost sensation, or suggest all robot AI fails at contact.
Benchmarks need to move with the field.
Failure rates from a 2025 paper using π0 cannot represent today’s general-purpose models. Following experience-based learning in the π0.6 family, Physical Intelligence announced π0.7 in April 2026, combining diverse data and conditional prompts for generalization. It also uses world models to generate visual subgoals.
The March 2026 RLT work connects VLA representations to a small reinforcement-learning policy to improve precision stages. This progress shows there is more than one route to solving contact tasks. Sensor value must be measured against current baselines with the same tasks, robots, and learning or adaptation budgets.
References
- Physical Intelligence · π0.7 ↗
April 16, 2026. Official announcement covering generalization, diverse prompts, and visual subgoals.
- Physical Intelligence · Precise Manipulation with Efficient Online RL ↗
March 19, 2026. Online adaptation of precision stages through RL tokens.
With feeling, the questions change.
Contact signals add “What am I caught on?” to “Where should I move?” Harness tension can prompt less pulling or a change of supporting hand. Surface work calls for coordinated pressure and feed speed. A seized part may require release and reapproach, rather than more pushing.
These are capabilities we intend to validate, not promised product performance. ContactBench defines variations and failures so we can evaluate whether sensing improves success, damage, recovery, and task time.
Touch is more than an image.
Feeding tactile images into a vision model is a useful starting point. We want to represent contact direction, magnitude, rate of change, and their relationship to action more directly. When force rises suddenly, what does a skilled person stop, and which way do they move next?
Alongside contact representations for VLAs, we consider world models that predict post-contact states and action outcomes. This direction combines simulated force data with real operator contact and response histories. World models are not all vision-only, and we do not claim our proposed system is complete.
It’s how force is used that matters.
A harder grip is not always a better grip. Skill means reading the part, applying only the force needed, and adapting when reality differs from expectation. That is why we record human contact intelligence: to help robots make better decisions with what they feel.