CoRL 2026 K UBM - Heungwoo/research GitHub Wiki
CoRL 2026 β K-UBM: Koopman Behavioral Models as Pseudo Planners for Visuo-Motor Dexterity
Venue: CoRL 2026 (Austin, TX, Nov 9β12) Β· Georgia Tech. Paper: arXiv 2602.07413. Representative of: prediction-as-execution-monitor β a structured latent (Koopman) system acts as a pseudo planner and its predicted-vs-observed visual flow triggers replanning. Companions: World Models Β· Dexterous Manipulation Β· CoRL 2026 survey.

1. Problem
Contemporary visuo-motor dexterity policies are reactive mappings that regress observations to fixed-horizon action chunks, treating manipulation as a sequence of independent decisions rather than a continuous physical process. This forces a trade-off between temporal coherence (long chunks) and responsiveness (short chunks), and such data/compute-heavy models remain brittle for multi-fingered manipulation under disturbance and occlusion.
2. Method
The authors propose Unified Behavioral Models (UBMs): represent a dexterous skill as a coupled dynamical system in which robot action and environmental visual features co-evolve. They instantiate this as Koopman-UBM (K-UBM). Robot joint commands and learned visual features are concatenated into a unified behavioral state ΞΎ_t = [a_t, Ο_t], which a spectral encoder lifts into a latent z_t whose joint evolution becomes linear: z_{t+1} = K z_t, with a single learnable Koopman matrix K.
- Pseudo planner (not chunking): instead of predicting a short action chunk, the linear system is rolled out from initial conditions to produce a full, flexible-horizon trajectory β recovering variable-length "dynamic chunking" without sequential decoding.
- Flow prediction as its own monitor: the same rollout predicts future visual features/flow. At runtime K-UBM compares predicted vs. observed visual flow; when a disturbance drives the discrepancy past a threshold it event-triggers a replan (re-initializes z_t and computes a new coherent trajectory), staying reactive without continuous feedback coupling.
3. Results
- 7 simulated tasks β DexArt (Bucket, Laptop, Toilet) + Adroit (Door, Tool use, In-hand reorientation, Relocation) β and 4 real dexterous tasks (Flower Arrangement, Chips Pouring, Whiteboard Erasing, Dinner Serving).
- Sim: K-UBM reaches the second-highest mean success rate with the most consistent performance across tasks (baselines: Diffusion Policy, ACT, UVA, KODex, KOROL), with little per-task architecture tuning.
- Real: 85.0% success (vs. ACT 92.5%) while cutting motion jerk sharply β arm jerk reduced ~66β9Γ and hand jerk ~33β4Γ vs. baselines.
- Inference: latent rollout is ~0.017 ms/step (Table 1: ~0.42β0.47 ms feature-encode), versus ~30β36 ms for Diffusion Policy.
4. Why it matters
K-UBM reframes a dexterous policy as a structured latent dynamical system rather than a black-box chunk predictor: linearity gives near-free, arbitrary-horizon rollouts, and predicting environmental visual flow turns the policy into its own runtime monitor for principled, event-triggered replanning β buying smooth execution, occlusion robustness, and orders-of-magnitude faster inference.
Limitations (reviewer): linear latent dynamics smooth over sharp contact discontinuities (high-frequency impacts may need switching/hybrid Koopman operators); open-loop pseudo-planning can fail in highly stochastic environments with unpredictable object dynamics; and simultaneous occlusion + disturbance can delay replanning and cause failure (would benefit from tactile feedback).
5. Links
- arXiv 2602.07413
- Survey: CoRL 2026 Β· Related: World Models Β· Dexterous Manipulation
β Back to CoRL 2026 survey Β· Home