CoRL 2026 K UBM - Heungwoo/research GitHub Wiki

CoRL 2026 β€” K-UBM: Koopman Behavioral Models as Pseudo Planners for Visuo-Motor Dexterity

Venue: CoRL 2026 (Austin, TX, Nov 9–12) Β· Georgia Tech. Paper: arXiv 2602.07413. Representative of: prediction-as-execution-monitor β€” a structured latent (Koopman) system acts as a pseudo planner and its predicted-vs-observed visual flow triggers replanning. Companions: World Models Β· Dexterous Manipulation Β· CoRL 2026 survey.

Unified Behavioral Models: robot action and environmental visual flow co-evolve in a shared Koopman latent (figure from Han et al., arXiv 2602.07413, Β© the authors)

1. Problem

Contemporary visuo-motor dexterity policies are reactive mappings that regress observations to fixed-horizon action chunks, treating manipulation as a sequence of independent decisions rather than a continuous physical process. This forces a trade-off between temporal coherence (long chunks) and responsiveness (short chunks), and such data/compute-heavy models remain brittle for multi-fingered manipulation under disturbance and occlusion.

2. Method

The authors propose Unified Behavioral Models (UBMs): represent a dexterous skill as a coupled dynamical system in which robot action and environmental visual features co-evolve. They instantiate this as Koopman-UBM (K-UBM). Robot joint commands and learned visual features are concatenated into a unified behavioral state ΞΎ_t = [a_t, Ο†_t], which a spectral encoder lifts into a latent z_t whose joint evolution becomes linear: z_{t+1} = K z_t, with a single learnable Koopman matrix K.

  • Pseudo planner (not chunking): instead of predicting a short action chunk, the linear system is rolled out from initial conditions to produce a full, flexible-horizon trajectory β€” recovering variable-length "dynamic chunking" without sequential decoding.
  • Flow prediction as its own monitor: the same rollout predicts future visual features/flow. At runtime K-UBM compares predicted vs. observed visual flow; when a disturbance drives the discrepancy past a threshold it event-triggers a replan (re-initializes z_t and computes a new coherent trajectory), staying reactive without continuous feedback coupling.

3. Results

  • 7 simulated tasks β€” DexArt (Bucket, Laptop, Toilet) + Adroit (Door, Tool use, In-hand reorientation, Relocation) β€” and 4 real dexterous tasks (Flower Arrangement, Chips Pouring, Whiteboard Erasing, Dinner Serving).
  • Sim: K-UBM reaches the second-highest mean success rate with the most consistent performance across tasks (baselines: Diffusion Policy, ACT, UVA, KODex, KOROL), with little per-task architecture tuning.
  • Real: 85.0% success (vs. ACT 92.5%) while cutting motion jerk sharply β€” arm jerk reduced ~66–9Γ— and hand jerk ~33–4Γ— vs. baselines.
  • Inference: latent rollout is ~0.017 ms/step (Table 1: ~0.42–0.47 ms feature-encode), versus ~30–36 ms for Diffusion Policy.

4. Why it matters

K-UBM reframes a dexterous policy as a structured latent dynamical system rather than a black-box chunk predictor: linearity gives near-free, arbitrary-horizon rollouts, and predicting environmental visual flow turns the policy into its own runtime monitor for principled, event-triggered replanning β€” buying smooth execution, occlusion robustness, and orders-of-magnitude faster inference.

Limitations (reviewer): linear latent dynamics smooth over sharp contact discontinuities (high-frequency impacts may need switching/hybrid Koopman operators); open-loop pseudo-planning can fail in highly stochastic environments with unpredictable object dynamics; and simultaneous occlusion + disturbance can delay replanning and cause failure (would benefit from tactile feedback).

5. Links

← Back to CoRL 2026 survey Β· Home