ICLR 2026 DexMove - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 Affiliation: ShanghaiTech · BIGAI · Beihang Category: Dexterous Manipulation — Non-Prehensile / Tactile / Wrist-Finger Synergy Trend tag: Tactile · contact-rich · flow-matching policy · wearable data
flowchart LR
subgraph SIM[Simulation pipeline — 2M sequences]
YCB[88 YCB objects → 352 instances<br/>random scale + rotation] --> Contact[Contact establishment<br/>uniform wrist pose sample + IK]
Contact --> Filter[Reject self-collision / penetration<br/>412k valid configurations]
Filter --> Reject[Rejection sampling:<br/>fingertips stable contact over 50 cm displacement]
Reject --> Force[Synthesize force-conditioned trajectories<br/>indentation depth as force surrogate<br/>Gaussian augmentation along contact normal]
end
subgraph REAL[Wearable tactile capture]
Wear[R-Tac fingertip sensors<br/>OV9281 mono camera 120 FPS<br/>33 markers per finger] --> H[20 objects, ~300k frames @ 30 FPS]
end
SIM --> EC[Establish-Contact FM policy<br/>PointNet++ + FiLM + flow matching]
H --> TaFo[TaFo-Net<br/>per-finger spatial enc → cross-finger attn →<br/>finger-wise causal temporal attn]
SIM --> POL[DexMove-Policy<br/>Transformer enc+dec flow matching<br/>past Tp=5, future Tf=5, 30 Hz]
TaFo --> POL
POL --> Run[Franka FR3 + Allegro Hand<br/>~22 ms inference / chunk]
Non-prehensile manipulation (pushing, sliding, pivoting without enclosure grasps) with multi-fingered dexterous hands is largely unexplored. Two specific blockers:
- Data scarcity. No large-scale dataset covers force-aware, multi-finger non-prehensile trajectories. Teleoperation suffers from missing haptic feedback (lower fidelity, lower success rates); pure simulation has soft-body and contact-modeling gaps; tactile gloves have layout mismatches between human hands and robot hands.
- No wrist-finger coordination policy. Existing dexterous manipulation focuses on grasping; pushing/pivoting work uses single-contact tools or grippers. Multi-fingered hands couple wrist and finger forces through hand-object dynamics, but no planner exists for this combined control problem.
The paper's argument: multi-fingered hands are intrinsically better for non-prehensile work because they can establish distributed contacts that are more stable than a single-contact rod or two-finger gripper — particularly for thin, cylindrical, or round objects where pushing dynamics are otherwise unpredictable.
Hand-object contact establishment. Uniformly sample wrist poses (R₀^wrist, T₀^wrist). For each fingertip, compute displacement d to the nearest object surface; perturb with Gaussian noise ε to produce d̂ = d + ε. Solve IK with:
Â₀^hand = argmin ‖FK(A₀^hand, R₀^wrist, T₀^wrist) − P₀^TIP‖² + w_pinch · L_region
where L_region = ‖d^TIP − d̂‖² encourages contacts inside the tactile sensor's effective region. Using 88 YCB objects × random scale & rotation = 352 instances; 1024-2048 candidates per instance → 412k valid contact configurations after collision filtering.
Force-conditioned trajectories. Repositioning controlled by (A^hand, R^wrist, T^wrist) inducing 3-DoF object motion (x, y, yaw). Instead of iLQR (Li & Todorov 2004), the authors use rejection sampling: in MuJoCo, translate the hand along random directions; accept if all fingertips maintain stable contact over 50 cm displacement.
For each accepted direction, uniformly sample object target (P_target^obj, ω_target^obj). Under the no-slip assumption:
P_t^tip = P_t^obj + R_z(ω_t^obj) (P_0^TIP − P_0^obj), for t = 0, ..., T.
Contact-force synthesis. Approximate normal force from indentation depth:
G ≈ D_sensor = r − distance(P_t^TIP, surface)
Augment by displacing each fingertip along the contact normal: P̂_t^TIP = P_t^TIP + n · N(0, σ). Recover joint and wrist configs via IK with a wrist-motion regulariser L^wrist so that the solution biases toward finger-driven rather than arm-driven manipulation. Filter trajectories that leave the workspace. Total: 2M sequences.
A wearable exoskeleton, isomorphic to the Allegro Hand, mounts R-Tac vision-based tactile sensors (Lin et al. 2025) on each human fingertip:
- Camera: OV9281 global-shutter monochrome, 160° FoV, 640×480 @ 120 FPS, fixed exposure.
- Illumination: 8× 4000K white LEDs (2835 package) in an annular PCB.
- Elastomer: PDMS base + Ecoflex 00-10 layer with 33 visual markers for shear-force detection.
- Tactile vector field: V ∈ ℝ^{v×4} where v = 33 (markers); channels are (normal force magnitude, shear direction-x, shear direction-y, shear magnitude).
- Marker tracking via Farneback optical flow (Farnebäck 2003); depth reconstruction via grayscale → indentation-depth LUT calibrated with a 2 mm spherical indenter.
- Data collected: 20 objects, ~300k frames @ 30 FPS.
The exoskeleton is isomorphic to the robot hand, deliberately minimising the domain gap.
4.1 Establish-Contact (Flow Matching).
- Input: object point cloud D, target pose (P^obj_target, ω^obj_target).
- Output: (A^hand_0, R^wrist_0, T^wrist_0).
- Backbone: PointNet++ for point-cloud features, FiLM conditioning into a 5-layer MLP (widths 128, 128, 512, 1024, 1024), flow-matching loss
L_contact = E[‖(X_1 − X_0) − u(X_t, t, cond)‖²]with X_t = (1−t)X_0 + tX_1. - Training: batch 128, AdamW, LR 1e-4, 1.3M steps.
- FM chosen over diffusion policy for faster training and inference.
4.2 DexMove-Policy (transformer flow matching).
- State at time t (past Tp = 5 frames):
(P^hand, A^hand, R^wrist, T^wrist, P^obj, ω^obj, C, G)_{−Tp:0}where P^hand ∈ ℝ^{J×3} joint positions, C ∈ ℝ^{F×3} per-finger contact positions in local frame, G ∈ ℝ^F per-finger pressing force. - Target: future hand state X_1 over Tf = 5 frames.
- Architecture: Past tokens + global target token (linear projection of (P_target, θ_target)) + continuous time token (Fourier features → MLP) → Transformer encoder produces memory M ∈ ℝ^{(Tp+2)×d}. The noised state X_t and planned future force G_{1:Tf} are FiLM-fused into query tokens for the Transformer decoder, which outputs the velocity field.
- Training: 200k steps, batch 2048; then 20k steps of ReFlow (Liu et al. 2022) to compress to a 10-step inference sampler.
- Inference: ~22 ms/chunk on RTX 4090, executed at 30 Hz.
4.3 TaFo-Net (tactile force planner). Given target pose, Tp = 5 past frames of object states, and per-finger tactile vector fields V_{−Tp:0} ∈ ℝ^{Tp×F×v×C}, predict future tactile fields V_{1:Tf} from which per-finger forces G_{1:Tf} are extracted. Three-stage architecture:
- Per-finger spatial encoding. Each V_{t,f} ∈ ℝ^{v×C} → token U_{t,f} via a lightweight transformer with learnable + geometry-informed marker positional embeddings.
-
Cross-finger attention. For each frame i, multi-head self-attention across F fingers augmented with per-finger type embeddings g_f:
Ũ_{i,1:F} = CF(U_{i,1:F} + g_{1:F}). - Finger-wise causal temporal attention. Causal mask so a query at i can only attend to tokens at times ≤ i, preventing future leakage.
Training loss L_rec = Σ_t Σ_f ‖V̂_{t,f} − V_{t,f}‖², with random dropout of time steps, fingers, and markers for robustness.
- Robot: Franka FR3 + Allegro Hand. Position control through ROS — Cartesian for the arm, joint-space PID for the hand.
- Vision: 3× Realsense D435i depth cameras (one near elbow, two on opposite sides). Object pose via ArUco markers (calibrated; markerless results also reported using FoundationPose).
- Compute: RTX 4090 deployment.
Six everyday objects: LEGO, mouse, keyboard, book, large can, small can. Two surfaces: Friction A (clean table) and Friction B (tape strips, unseen during data collection). Initial-yaw-error bins:
| Method | 0° < ω_target < 30° (A / B) | 30°–60° (A / B) | 60°–90° (A / B) |
|---|---|---|---|
| Open-loop | 36.7 / 10.0 | 23.3 / 0.0 | 3.3 / 0.0 |
| DyWA (Lyu et al. 2025) | 50.0 / 36.7 | 46.7 / 30.0 | 50.0 / 33.3 |
| CORN (Cho et al. 2024) | 43.3 / 36.7 | 46.7 / 40.0 | 43.3 / 43.3 |
| DexMove | 86.7 / 86.7 | 80.0 / 83.3 | 70.0 / 60.0 |
DexMove maintains performance under unseen friction (gap is small, A vs. B); DyWA and CORN show pronounced degradation, reflecting their sensitivity to spatial friction variability.
| Method | 0–15 cm | 15–30 cm | 30–45 cm |
|---|---|---|---|
| DyWA | 36.1 | 52.2 | 60.6 |
| CORN | 41.4 | 54.5 | 62.1 |
| DexMove | 8.3 | 10.9 | 12.4 |
DexMove is 3-5× faster because multi-finger contact reduces the number of action primitives needed to reach a target pose.
Across the six objects, average success 77.8%; +36.6% over ablated baselines, ~300% efficiency improvement over baseline methods (i.e., ~4× speed).
| Object | Trials | Success |
|---|---|---|
| Rag doll | 30 | 96.7% |
| Tissue packet | 30 | 100% |
Deformability actually helps — compliant contact stabilises contact formation.
Random stacking of objects beneath the manipulated items. Tested on book, large can, LEGO with two conditions (w/o finetune, w/ finetune on 15 min of uneven-surface tactile data + masked contact intervals). Specific numbers not stated in extracted body but the paper claims robustness with light finetuning.
Replacing ArUco markers with FoundationPose: success rates of 16.7, 13.3, 93.3, 96.7, 60.0, 76.6% for the six objects. Hand-occlusion-induced pose estimation errors hurt small objects the most.
| Noise σ | Err-MSE (book) | SR (book) | Err-MSE (can) | SR (can) |
|---|---|---|---|---|
| 0 | 0.0112 | 90.0% | 0.0351 | 63.3% |
| 0.05 | 0.0108 | 86.7% | 0.0615 | 56.7% |
| 0.1 | 0.0415 | 80.0% | 0.1239 | 43.3% |
| 0.2 | 0.1721 | 53.3% | 0.3005 | 20.0% |
| 0.4 | 0.3219 | 13.3% | 0.5312 | 3.3% |
Robust up to σ = 0.1 (sensor noise + TaFo-Net prediction errors). Attributed to (i) noisy training tactile signals, (ii) random dropout during training.
The learned policy generalizes to language-conditioned, long-horizon tasks — one of the paper's three stated significance pillars. Three demonstrated scenarios: (i) structured sorting ("move box A to region 1"); (ii) language-driven human–machine collaboration, where a VLM (SoFar, Qi et al. 2025) converts a natural-language command (e.g., "put the grip of the electric drill into a person's hand") into a 3-DoF target pose fed to the policy for a non-prehensile handover; (iii) desktop tidying, relocating each item to an assigned position from a predefined layout.
| Method | LEGO | Mouse | Book | Keyboard | Large Can | Small Can |
|---|---|---|---|---|---|---|
| Wrist-Only (auto, fingers locked) | 13.3 | 0.0 | 33.3 | 20.0 | 0.0 | 0.0 |
| Wrist-Only* (teleoperated) | 0.0 | 73.3 | 100.0 | 100.0 | 6.7 | 10.0 |
| w/o Cross-Finger | 13.3 | 3.3 | 63.3 | 50.0 | 0.0 | 3.3 |
| w/o Shear-Force | 70.0 | 66.7 | 33.3 | 13.3 | 0.0 | 0.0 |
| w Heuristic Force | 36.7 | 43.3 | 66.7 | 0.0 | 0.0 | 0.0 |
| DexMove (Ours) | 66.7 | 86.7 | 90.0 | 90.0 | 63.3 | 70.0 |
Findings:
- Wrist-only can handle flat/planar objects (book, keyboard) when teleoperated — confirming the wrist contributes substantially — but rarely succeeds on objects requiring non-coplanar fingertip contacts (LEGO, mouse) or shape-induced grasp adjustments (cans).
- Cross-finger attention is essential — without it TaFo-Net fails on round/heavy objects (large can goes from 63.3 → 0%).
- Shear force is critical for slip detection and heavier objects. Without it, the model converges on smoothed averaged states — fine for light objects (LEGO, mouse) but fails on cans.
- Heuristic force (slip-detection + force increment, à la Lin et al. 2025) cannot replace TaFo-Net — incrementing force after slip is reactive, not predictive.
The authors explicitly list:
- Articulated objects. Objects with movable parts (e.g., a telephone with a handset) shift during manipulation and destabilise contact.
- Spherical objects. Tend to roll; stable initial contact is difficult; slippage risk increases.
- Hand-pose failure modes. Some grasps cause the object to topple — e.g., pushing a tall can while gripping only its lid causes a tip-over and restricts rotational motion.
Additional implicit limitations from experiments:
- Markerless pose estimation degrades small-object performance (16.7-13.3% on LEGO and mouse vs. 90% with ArUco).
- The simulation-to-real bridge depends on the isomorphic exoskeleton; non-isomorphic robot hands would require new data collection.
- Friction generalisation tested in only two regimes (clean vs. tape).
- Authors plan to extend to prehensile + non-prehensile integration in future work.
DexMove opens the non-prehensile regime for multi-fingered hands by:
- Treating wrist and fingers as a coupled control problem rather than two separate stages — the wrist-motion regulariser in IK biases trajectories toward finger-driven manipulation.
- Bringing real tactile data into the loop via the isomorphic wearable with R-Tac sensors, sidestepping the layout mismatch of tactile gloves.
- Using flow matching (faster than diffusion policy) for both the contact policy and the trajectory policy, with ReFlow compression to 10 inference steps.
The skill also plugs into language-conditioned, long-horizon pipelines (object sorting, VLM-driven handover via SoFar, desktop tidying — §5.5), showing the non-prehensile primitive composes into higher-level task planning.
Versus related work:
- CORN (Cho et al. 2024) & DyWA (Lyu et al. 2025): gripper-based non-prehensile baselines that lose accuracy on cylindrical objects (single contact point). DexMove leverages distributed multi-finger contacts.
- DemoGrasp / DexNDM: grasp-centric dexterous learning. Complementary — DexMove handles the non-grasp regime they exclude.
- DexUMI: wearable dexterous teleoperation for grasping. DexMove extends the isomorphic-wearable idea to force-aware non-prehensile data.
- 3D-ViTac (Huang et al. 2024), Higuera et al. 2025: tactile-rich manipulation. DexMove adds the multi-contact non-prehensile angle and the wrist-finger synergy claim.
- DexNoMa (Li et al. 2025b): dexterous non-prehensile manipulation, but on smaller object sets and without tactile sensing.
The 77.8% real-world average across six varied objects, with 3-5× faster execution than gripper baselines and 96-100% on deformable objects, is the strongest quantitative argument for why multi-fingered hands matter for non-prehensile tasks — they are not just more dexterous, they are more stable and more efficient than the gripper alternatives.
- OpenReview: https://openreview.net/forum?id=dT3ZciXvNX
- Project: https://peilin-666.github.io/projects/DexMove/
← Back to ICLR-2026