Tactile-VLA: connecting object knowledge to touch
Between knowing what an object is and deciding how to hold it.
INTERACT ROBOTICS · Research Notes · Reviewed September 7, 2026
Different objects. Different ways to hold on.
A heavy, rigid object and a light, fragile one call for different contact strategies. Tactile-VLA connects tactile observations with a VLA’s visual and language representations to bring knowledge of physical properties into manipulation.
Insertion and separation comparisons within the paper
| Model | USB | Charger |
|---|---|---|
| π0-base | 5% | 40% |
| π0-fast | 0% | 25% |
| Tactile-VLA | 35% | 90% |
Original paper, Table 1. The 40% baseline compared with 90% on the charger task is π0-base. Do not substitute the π0-fast result.
Reading generalization carefully.
Choosing a contact strategy for an unfamiliar object is a promising direction. The charger result cannot be generalized to every insertion task; remaining USB failures matter too. Knowledge of an object category and observation of actual contact have different roles.
The questions we bring to it.
Which tactile components help? We want to separate the contributions of contact presence, distribution, slip, and force direction to action choice. Evaluate clear and occluded vision, new materials, and new shapes separately.
References
- Tactile-VLA: Unlocking Vision-Language-Action Model’s Physical Knowledge for Tactile Generalization ↗
2025 / Table 1. Evaluated against the π0 baselines of that time; not evidence of limitations in π0.7.