ICLR 2026 DexMove - Heungwoo/research GitHub Wiki

DexMove — Tactile-Guided Non-Prehensile Dexterous Manipulation

Venue: ICLR 2026 Affiliation: ShanghaiTech · BIGAI · Beihang Category: Dexterous Manipulation — Non-Prehensile / Tactile / Wrist-Finger Synergy Trend tag: Tactile · contact-rich · flow-matching policy · wearable data

Approach diagram

flowchart LR
  subgraph SIM[Simulation pipeline — 2M sequences]
    YCB[88 YCB objects → 352 instances<br/>random scale + rotation] --> Contact[Contact establishment<br/>uniform wrist pose sample + IK]
    Contact --> Filter[Reject self-collision / penetration<br/>412k valid configurations]
    Filter --> Reject[Rejection sampling:<br/>fingertips stable contact over 50 cm displacement]
    Reject --> Force[Synthesize force-conditioned trajectories<br/>indentation depth as force surrogate<br/>Gaussian augmentation along contact normal]
  end
  subgraph REAL[Wearable tactile capture]
    Wear[R-Tac fingertip sensors<br/>OV9281 mono camera 120 FPS<br/>33 markers per finger] --> H[20 objects, ~300k frames @ 30 FPS]
  end
  SIM --> EC[Establish-Contact FM policy<br/>PointNet++ + FiLM + flow matching]
  H --> TaFo[TaFo-Net<br/>per-finger spatial enc → cross-finger attn →<br/>finger-wise causal temporal attn]
  SIM --> POL[DexMove-Policy<br/>Transformer enc+dec flow matching<br/>past Tp=5, future Tf=5, 30 Hz]
  TaFo --> POL
  POL --> Run[Franka FR3 + Allegro Hand<br/>~22 ms inference / chunk]
Loading

Problem

Non-prehensile manipulation (pushing, sliding, pivoting without enclosure grasps) with multi-fingered dexterous hands is largely unexplored. Two specific blockers:

  1. Data scarcity. No large-scale dataset covers force-aware, multi-finger non-prehensile trajectories. Teleoperation suffers from missing haptic feedback (lower fidelity, lower success rates); pure simulation has soft-body and contact-modeling gaps; tactile gloves have layout mismatches between human hands and robot hands.
  2. No wrist-finger coordination policy. Existing dexterous manipulation focuses on grasping; pushing/pivoting work uses single-contact tools or grippers. Multi-fingered hands couple wrist and finger forces through hand-object dynamics, but no planner exists for this combined control problem.

The paper's argument: multi-fingered hands are intrinsically better for non-prehensile work because they can establish distributed contacts that are more stable than a single-contact rod or two-finger gripper — particularly for thin, cylindrical, or round objects where pushing dynamics are otherwise unpredictable.

Method (detailed)

3.1 Trajectory synthesis (2M sequences)

Hand-object contact establishment. Uniformly sample wrist poses (R₀^wrist, T₀^wrist). For each fingertip, compute displacement d to the nearest object surface; perturb with Gaussian noise ε to produce d̂ = d + ε. Solve IK with:

Â₀^hand = argmin ‖FK(A₀^hand, R₀^wrist, T₀^wrist) − P₀^TIP‖² + w_pinch · L_region

where L_region = ‖d^TIP − d̂‖² encourages contacts inside the tactile sensor's effective region. Using 88 YCB objects × random scale & rotation = 352 instances; 1024-2048 candidates per instance → 412k valid contact configurations after collision filtering.

Force-conditioned trajectories. Repositioning controlled by (A^hand, R^wrist, T^wrist) inducing 3-DoF object motion (x, y, yaw). Instead of iLQR (Li & Todorov 2004), the authors use rejection sampling: in MuJoCo, translate the hand along random directions; accept if all fingertips maintain stable contact over 50 cm displacement.

For each accepted direction, uniformly sample object target (P_target^obj, ω_target^obj). Under the no-slip assumption:

P_t^tip = P_t^obj + R_z(ω_t^obj) (P_0^TIP − P_0^obj), for t = 0, ..., T.

Contact-force synthesis. Approximate normal force from indentation depth:

G ≈ D_sensor = r − distance(P_t^TIP, surface)

Augment by displacing each fingertip along the contact normal: P̂_t^TIP = P_t^TIP + n · N(0, σ). Recover joint and wrist configs via IK with a wrist-motion regulariser L^wrist so that the solution biases toward finger-driven rather than arm-driven manipulation. Filter trajectories that leave the workspace. Total: 2M sequences.

3.2 Wearable tactile capture

A wearable exoskeleton, isomorphic to the Allegro Hand, mounts R-Tac vision-based tactile sensors (Lin et al. 2025) on each human fingertip:

  • Camera: OV9281 global-shutter monochrome, 160° FoV, 640×480 @ 120 FPS, fixed exposure.
  • Illumination: 8× 4000K white LEDs (2835 package) in an annular PCB.
  • Elastomer: PDMS base + Ecoflex 00-10 layer with 33 visual markers for shear-force detection.
  • Tactile vector field: V ∈ ℝ^{v×4} where v = 33 (markers); channels are (normal force magnitude, shear direction-x, shear direction-y, shear magnitude).
  • Marker tracking via Farneback optical flow (Farnebäck 2003); depth reconstruction via grayscale → indentation-depth LUT calibrated with a 2 mm spherical indenter.
  • Data collected: 20 objects, ~300k frames @ 30 FPS.

The exoskeleton is isomorphic to the robot hand, deliberately minimising the domain gap.

4. Policies

4.1 Establish-Contact (Flow Matching).

  • Input: object point cloud D, target pose (P^obj_target, ω^obj_target).
  • Output: (A^hand_0, R^wrist_0, T^wrist_0).
  • Backbone: PointNet++ for point-cloud features, FiLM conditioning into a 5-layer MLP (widths 128, 128, 512, 1024, 1024), flow-matching loss L_contact = E[‖(X_1 − X_0) − u(X_t, t, cond)‖²] with X_t = (1−t)X_0 + tX_1.
  • Training: batch 128, AdamW, LR 1e-4, 1.3M steps.
  • FM chosen over diffusion policy for faster training and inference.

4.2 DexMove-Policy (transformer flow matching).

  • State at time t (past Tp = 5 frames): (P^hand, A^hand, R^wrist, T^wrist, P^obj, ω^obj, C, G)_{−Tp:0} where P^hand ∈ ℝ^{J×3} joint positions, C ∈ ℝ^{F×3} per-finger contact positions in local frame, G ∈ ℝ^F per-finger pressing force.
  • Target: future hand state X_1 over Tf = 5 frames.
  • Architecture: Past tokens + global target token (linear projection of (P_target, θ_target)) + continuous time token (Fourier features → MLP) → Transformer encoder produces memory M ∈ ℝ^{(Tp+2)×d}. The noised state X_t and planned future force G_{1:Tf} are FiLM-fused into query tokens for the Transformer decoder, which outputs the velocity field.
  • Training: 200k steps, batch 2048; then 20k steps of ReFlow (Liu et al. 2022) to compress to a 10-step inference sampler.
  • Inference: ~22 ms/chunk on RTX 4090, executed at 30 Hz.

4.3 TaFo-Net (tactile force planner). Given target pose, Tp = 5 past frames of object states, and per-finger tactile vector fields V_{−Tp:0} ∈ ℝ^{Tp×F×v×C}, predict future tactile fields V_{1:Tf} from which per-finger forces G_{1:Tf} are extracted. Three-stage architecture:

  1. Per-finger spatial encoding. Each V_{t,f} ∈ ℝ^{v×C} → token U_{t,f} via a lightweight transformer with learnable + geometry-informed marker positional embeddings.
  2. Cross-finger attention. For each frame i, multi-head self-attention across F fingers augmented with per-finger type embeddings g_f: Ũ_{i,1:F} = CF(U_{i,1:F} + g_{1:F}).
  3. Finger-wise causal temporal attention. Causal mask so a query at i can only attend to tokens at times ≤ i, preventing future leakage.

Training loss L_rec = Σ_t Σ_f ‖V̂_{t,f} − V_{t,f}‖², with random dropout of time steps, fingers, and markers for robustness.

Hardware setup

  • Robot: Franka FR3 + Allegro Hand. Position control through ROS — Cartesian for the arm, joint-space PID for the hand.
  • Vision: 3× Realsense D435i depth cameras (one near elbow, two on opposite sides). Object pose via ArUco markers (calibrated; markerless results also reported using FoundationPose).
  • Compute: RTX 4090 deployment.

Comprehensive Results

Main benchmark (Table 1, success rate %, 30 trials per setting)

Six everyday objects: LEGO, mouse, keyboard, book, large can, small can. Two surfaces: Friction A (clean table) and Friction B (tape strips, unseen during data collection). Initial-yaw-error bins:

Method 0° < ω_target < 30° (A / B) 30°–60° (A / B) 60°–90° (A / B)
Open-loop 36.7 / 10.0 23.3 / 0.0 3.3 / 0.0
DyWA (Lyu et al. 2025) 50.0 / 36.7 46.7 / 30.0 50.0 / 33.3
CORN (Cho et al. 2024) 43.3 / 36.7 46.7 / 40.0 43.3 / 43.3
DexMove 86.7 / 86.7 80.0 / 83.3 70.0 / 60.0

DexMove maintains performance under unseen friction (gap is small, A vs. B); DyWA and CORN show pronounced degradation, reflecting their sensitivity to spatial friction variability.

Execution efficiency (Table 2, seconds to completion)

Method 0–15 cm 15–30 cm 30–45 cm
DyWA 36.1 52.2 60.6
CORN 41.4 54.5 62.1
DexMove 8.3 10.9 12.4

DexMove is 3-5× faster because multi-finger contact reduces the number of action primitives needed to reach a target pose.

Aggregate metric (paper abstract)

Across the six objects, average success 77.8%; +36.6% over ablated baselines, ~300% efficiency improvement over baseline methods (i.e., ~4× speed).

Deformable objects (§5.3)

Object Trials Success
Rag doll 30 96.7%
Tissue packet 30 100%

Deformability actually helps — compliant contact stabilises contact formation.

Uneven surfaces (Fig. 6)

Random stacking of objects beneath the manipulated items. Tested on book, large can, LEGO with two conditions (w/o finetune, w/ finetune on 15 min of uneven-surface tactile data + masked contact intervals). Specific numbers not stated in extracted body but the paper claims robustness with light finetuning.

Markerless pose estimation (FoundationPose, §5.3)

Replacing ArUco markers with FoundationPose: success rates of 16.7, 13.3, 93.3, 96.7, 60.0, 76.6% for the six objects. Hand-occlusion-induced pose estimation errors hurt small objects the most.

Tactile noise robustness (Table 4)

Noise σ Err-MSE (book) SR (book) Err-MSE (can) SR (can)
0 0.0112 90.0% 0.0351 63.3%
0.05 0.0108 86.7% 0.0615 56.7%
0.1 0.0415 80.0% 0.1239 43.3%
0.2 0.1721 53.3% 0.3005 20.0%
0.4 0.3219 13.3% 0.5312 3.3%

Robust up to σ = 0.1 (sensor noise + TaFo-Net prediction errors). Attributed to (i) noisy training tactile signals, (ii) random dropout during training.

Downstream applications (§5.5)

The learned policy generalizes to language-conditioned, long-horizon tasks — one of the paper's three stated significance pillars. Three demonstrated scenarios: (i) structured sorting ("move box A to region 1"); (ii) language-driven human–machine collaboration, where a VLM (SoFar, Qi et al. 2025) converts a natural-language command (e.g., "put the grip of the electric drill into a person's hand") into a 3-DoF target pose fed to the policy for a non-prehensile handover; (iii) desktop tidying, relocating each item to an assigned position from a predefined layout.

Ablation Studies (Table 3, success rate % across 6 objects)

Method LEGO Mouse Book Keyboard Large Can Small Can
Wrist-Only (auto, fingers locked) 13.3 0.0 33.3 20.0 0.0 0.0
Wrist-Only* (teleoperated) 0.0 73.3 100.0 100.0 6.7 10.0
w/o Cross-Finger 13.3 3.3 63.3 50.0 0.0 3.3
w/o Shear-Force 70.0 66.7 33.3 13.3 0.0 0.0
w Heuristic Force 36.7 43.3 66.7 0.0 0.0 0.0
DexMove (Ours) 66.7 86.7 90.0 90.0 63.3 70.0

Findings:

  1. Wrist-only can handle flat/planar objects (book, keyboard) when teleoperated — confirming the wrist contributes substantially — but rarely succeeds on objects requiring non-coplanar fingertip contacts (LEGO, mouse) or shape-induced grasp adjustments (cans).
  2. Cross-finger attention is essential — without it TaFo-Net fails on round/heavy objects (large can goes from 63.3 → 0%).
  3. Shear force is critical for slip detection and heavier objects. Without it, the model converges on smoothed averaged states — fine for light objects (LEGO, mouse) but fails on cans.
  4. Heuristic force (slip-detection + force increment, à la Lin et al. 2025) cannot replace TaFo-Net — incrementing force after slip is reactive, not predictive.

Limitations (stated by authors, §6)

The authors explicitly list:

  1. Articulated objects. Objects with movable parts (e.g., a telephone with a handset) shift during manipulation and destabilise contact.
  2. Spherical objects. Tend to roll; stable initial contact is difficult; slippage risk increases.
  3. Hand-pose failure modes. Some grasps cause the object to topple — e.g., pushing a tall can while gripping only its lid causes a tip-over and restricts rotational motion.

Additional implicit limitations from experiments:

  • Markerless pose estimation degrades small-object performance (16.7-13.3% on LEGO and mouse vs. 90% with ArUco).
  • The simulation-to-real bridge depends on the isomorphic exoskeleton; non-isomorphic robot hands would require new data collection.
  • Friction generalisation tested in only two regimes (clean vs. tape).
  • Authors plan to extend to prehensile + non-prehensile integration in future work.

Significance & Positioning

DexMove opens the non-prehensile regime for multi-fingered hands by:

  1. Treating wrist and fingers as a coupled control problem rather than two separate stages — the wrist-motion regulariser in IK biases trajectories toward finger-driven manipulation.
  2. Bringing real tactile data into the loop via the isomorphic wearable with R-Tac sensors, sidestepping the layout mismatch of tactile gloves.
  3. Using flow matching (faster than diffusion policy) for both the contact policy and the trajectory policy, with ReFlow compression to 10 inference steps.

The skill also plugs into language-conditioned, long-horizon pipelines (object sorting, VLM-driven handover via SoFar, desktop tidying — §5.5), showing the non-prehensile primitive composes into higher-level task planning.

Versus related work:

  • CORN (Cho et al. 2024) & DyWA (Lyu et al. 2025): gripper-based non-prehensile baselines that lose accuracy on cylindrical objects (single contact point). DexMove leverages distributed multi-finger contacts.
  • DemoGrasp / DexNDM: grasp-centric dexterous learning. Complementary — DexMove handles the non-grasp regime they exclude.
  • DexUMI: wearable dexterous teleoperation for grasping. DexMove extends the isomorphic-wearable idea to force-aware non-prehensile data.
  • 3D-ViTac (Huang et al. 2024), Higuera et al. 2025: tactile-rich manipulation. DexMove adds the multi-contact non-prehensile angle and the wrist-finger synergy claim.
  • DexNoMa (Li et al. 2025b): dexterous non-prehensile manipulation, but on smaller object sets and without tactile sensing.

The 77.8% real-world average across six varied objects, with 3-5× faster execution than gripper baselines and 96-100% on deformable objects, is the strongest quantitative argument for why multi-fingered hands matter for non-prehensile tasks — they are not just more dexterous, they are more stable and more efficient than the gripper alternatives.

Links

Related pages

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️