UniDex-ViTac: Learning Unified Visuo-Tactile Dexterous Manipulation Policy from Human Video Data

ICRA 2026 Workshop
Beyond Teleoperation: Learning from Diverse Human and Simulation Data
🎤 Oral Presentation

Hyesung Lee1,2 Si-Hwan Heo1 Sungwook Yang1
1KIST    2KAIST

Abstract

Learning dexterous robotic manipulation directly from human videos is fundamentally challenged by the kinematic embodiment gap and the lack of contact information that is unobservable in videos. To address these limitations, we present UniDex-ViTac, a unified visuo-tactile imitation learning framework that distills physically feasible, contact-rich trajectories generated by residual RL specialists into a single multi-task generalist policy. Crucially, our generalist operates on an expressive visuo-tactile representation that explicitly fuses global 3D point clouds with local binary tactile feedback. By effectively reasoning over both spatial geometry and local contact events, UniDex-ViTac achieves a 68.3% success rate in simulation and demonstrates robust Sim2Real transfer on a physical 16-DoF hand, achieving a 66.4% average success rate across diverse seen and unseen objects.

Overview

Pipeline Overview

Our pipeline reconstructs HOI trajectories from human videos, trains per-object Residual RL specialists to produce physically feasible robot motions, then distills 10,000 successful rollouts into a single multi-task Visuo-Tactile Generalist policy.

Visuo-Tactile Observation

Point Cloud One-Hot Encoding

Each point in the 512-point cloud carries a one-hot label: (1,0,0) camera points, (0,1,0) non-contacting fingertips, (0,0,1) contacting fingertips. This is combined with a 42-D proprioceptive state and a 4-D binary contact signal, enabling the policy to jointly reason over scene geometry and local contact events.

Visuo-Tactile Generalist Policy Architecture

Generalist Policy Architecture

Built on ACT with a CVAE encoder, the policy Transformer takes four tokens — point cloud (PointNet++), proprioception, binary contact, and a style latent — and predicts a 30-step action chunk of wrist pose and fingertip targets, executed via DLS arm IK and analytical hand IK.

Simulation Results

Simulation Results

(a) Our visuo-tactile policy achieves 68.3% average success on 10 YCB objects, outperforming ACT (53.4%), Diffusion Policy (45.2%), and BC-Transformer (20.2%). (b) Point cloud and binary contact are complementary — fusing both outperforms either alone (PCD Only: 55.5%, Contact Only: 58.4%). (c) Our dual fusion (onehot label + contact token) consistently performs best across all objects.

Qualitative Rollouts (videos play at 1× speed)

Real-World Results

Deployed on a Franka Emika Panda with a 16-DoF dexterous hand, our PC + Tactile policy achieves 66.4% overall success across seen and unseen objects, outperforming the PC Only baseline (54.5%).

Seen Objects (videos play at 1× speed)

Unseen Objects (videos play at 1× speed)

Category Object PC Only PC + Tac
Seen bleach_cleanser10/109/10
mustard_bottle7/107/10
power_drill7/108/10
potted_meat_can4/107/10
tomato_soup_can3/105/10
pudding_box2/106/10
Average55.0%70.0%
Unseen maxwell_coffee_can 9/10 9/10
brown_salt_box2/104/10
plastic_wine_cup7/106/10
blue_mug0/102/10
pringles_bottle9/1010/10
Average54.0%62.0%
Overall Average 54.5% 66.4%

Table 1: Real-world success rates (out of 10 trials) comparing the Point Cloud Only policy and the Point Cloud + Binary Tactile policy.

Citation

@inproceedings{unidex-vitac2026,
  title     = {UniDex-ViTac: Learning Unified Visuo-Tactile Dexterous Manipulation Policy from Human Video Data},
  author    = {Hyesung Lee and Si-Hwan Heo and Sungwook Yang},
  booktitle = {ICRA 2026 Workshop: Beyond Teleoperation},
  year      = {2026}
}