ICLR 2026 UniHM - Heungwoo/research GitHub Wiki

UniHM — Unified Dexterous Hand Manipulation with VLM

Venue: ICLR 2026 Category: Dexterous Manipulation Trend tag: Trend 6

arXiv: 2603.00732 (v1, 28 Feb 2026) Authors: Zhenhao Zhang, Jiaxin Liu, Ye Shi, Jingya Wang (ShanghaiTech University; InstAdapt)

Approach diagram

flowchart LR
  Inst[Open-vocabulary instruction] --> VLM[VLM: Qwen3-0.6B]
  Img[RGB-D input] --> VLM
  VLM --> Tok[Manipulation token sequence]
  H1[Shadow / Allegro / SVH / Leap ...] --> Enc[Per-hand encoders<br/>shared codebook K=8192]
  Enc --> Tokz[Unified Hand-Dexterous Tokenizer]
  Tokz --> VLM
  Tok --> Dec[Hand-specific decoders]
  Dec --> Phys[Physics-aware decoding /<br/>energy-based refinement]
  Phys --> Traj[Dexterous-hand trajectory]
Loading

Problem

Dexterous manipulation has historically used per-task, per-hand pipelines, and prior language-conditioned work largely targets static grasps rather than dynamic, multi-step interaction. A unified, language-conditioned framework that generates physically-feasible manipulation sequences across heterogeneous hand morphologies — without relying on real-world teleoperation data — had not previously existed.

Method

UniHM is a three-stage motion-synthesis pipeline (it generates manipulation trajectories, not closed-loop control):

  1. Morphology-agnostic motion tokenization — a VQ-VAE with per-hand encoders/decoders mapping heterogeneous dexterous hands into a single shared codebook (K=8192); new morphologies are aligned to a reference encoder via knowledge distillation (the Unified Hand-Dexterous Tokenizer).
  2. Language-guided generation — a VLM (Qwen3-0.6B, chosen for data-efficient convergence on scarce HOI data) fuses text, RGB-D perception, and token history to autoregressively produce manipulation token sequences.
  3. Physics-aware decoding — energy-based / segment-wise refinement enforces contact feasibility, smoothness, and temporal coherence.

Trained solely on closed-set HOI (hand-object interaction) datasets (DexYCB, OakInk), eliminating reliance on real-world teleoperation. Cross-hand sharing is via the shared codebook (related in spirit to OmniSAT).

Results

Evaluated on DexYCB / OakInk with motion-quality metrics (MPJPE, FID, finger-orientation/position losses) and real-world execution on multiple robot hands (Shadow, Allegro, SVH, Leap, Panda). Reported real-world seen-object success rates: ~65% Grab, 50% Pick&Place, 60% Pull&Push, 55% Open&Close. Generalizes across hand morphologies via the shared codebook plus distillation; the framework is positioned for open-world generalization despite training only on closed-set HOI data.

Significance

The first unified, language-conditioned framework for dynamic dexterous hand manipulation across heterogeneous morphologies, trained without teleoperation data. Note: this is a trajectory/motion-synthesis system (it plans physically-feasible sequences), not an end-to-end closed-loop visuomotor controller — the paper explicitly distinguishes itself from generalist control policies such as π0.

Links

Related pages

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️