ICML 2026 DLO Lab - Heungwoo/research GitHub Wiki
DLO-Lab — A Differentiable Simulator and Benchmark for Deformable Linear Object Manipulation
Venue: ICML 2026 (Poster) Category: Benchmark Affiliations: Junyi Cao, Yian Wang, Ziyan Xiong, Chunru Lin, Zhehuan Chen, Chuang Gan

Problem
Robotic manipulation of deformable linear objects (DLOs) — ropes, cables, wires, rubber bands — has long resisted general solutions. Prior work is narrow and task-specific, leaning on real-world demonstrations or handcrafted heuristics that do not scale to the wide variety of materials and tasks seen in practice, and collecting diverse real data is impractical. Existing simulators only support a subset of the material behaviors needed for generalizable DLO manipulation. DLOs are uniquely hard because: (1) their co-dimensional (essentially 1-D) nature demands high precision, (2) they have complex dynamics and varied topologies that are difficult to model, and (3) the topological complexity and grasp sensitivity make many tasks kinematically unsolvable under a poor grasp.
Method
DLO-Lab is a differentiable simulator explicitly designed for versatile DLO manipulation. It models a broad range of material properties — (in)extensibility, elasticity, bending plasticity, and complex interactions with other objects — giving a single foundation for both learning and evaluation.
- DLO modeling. Following the classic Discrete Elastic Rod (DER) formulation, a DLO is represented as a centerline of N_v vertices plus a set of adapted orthonormal frames on each edge, with internal behavior driven by a potential-energy function over vertex and frame states. Implementation details cover inextensibility constraints, bending plasticity, loop topology, and self-collision/friction.
- Coupling. A two-way coupling scheme connects DLOs to rigid bodies (using the Genesis rigid solver, with each rigid geometry as a time-varying signed distance field) and to soft bodies. Contact is handled with a soft-coupling exponential influence factor rather than a hard binary contact, which keeps the simulation differentiable.
- Gradient back-propagation. Because the whole physics stack is differentiable, analytic gradients flow through the physics timesteps — exploited both by first-order model-based RL and by system identification (see Results).
- DLO Agent. To inject structural priors, a specialized agent (1) uses a Vision-Language Model to propose strategic grasp points guided by physical commonsense (three proposal modes: candidate, coefficient, marker), avoiding the prohibitively large random-sampling search space, and (2) decomposes long-horizon tasks to maximize control authority.
The benchmark comprises 10 manipulation tasks (8 fixed-horizon, 2 long-horizon, e.g., Unknotting, Wiring-ring), each cast as a finite-horizon MDP with a state space over DLO vertices/velocities and robot DoF.
Results
The authors benchmark a spectrum of policy-learning methods: model-free RL (SAC, PPO), first-order model-based RL (SHAC, SAPO), and trajectory optimization (CMA-ES, gradient descent).
- Differentiability pays off in contact-rich tasks. On the precise topological Unknotting task, the FO-MBRL methods SHAC and SAPO reach an episodic return of ≈46, while PPO and SAC fail to progress (returns ≈3).
- Trajectory optimization is strong in the sample-constrained regime. Under the same sampling budget, MFRL generally underperforms trajectory optimization, which solves a simpler open-loop problem; even FO-MBRL lags sample-based CMA-ES on highly non-smooth tasks.
- Sim-to-real. Using the differentiable simulator, system identification calibrates real ropes' stretching and bending stiffness by minimizing the pixel-wise projection error between simulated and real rope masks, with gradients back-propagated through physics timesteps. Real-world deployment includes a closed-loop Wiring-ring evaluation reported over 12 trials.
Significance
DLO-Lab unifies the two properties prior DLO simulators lacked simultaneously — coupling with other materials and differentiability — into one platform, and pairs it with a representative benchmark and a VLM-guided agent. This gives the community a scalable, reproducible foundation for studying generalizable deformable-object manipulation, including data generation and VLA fine-tuning, without the bottleneck of collecting diverse real-world rope data.
Links
- arXiv: 2606.04206
- ICML 2026: https://icml.cc/virtual/2026/poster/63391
← Back to ICML-2026