ICML 2026 ManiSoft - Heungwoo/research GitHub Wiki
ManiSoft — A vision-language manipulation benchmark for soft continuum robots
Venue: ICML 2026 (Poster) Category: Benchmark Affiliations: Ziyu Wei, Luting Wang, Chen Gao, Li Wen, Si Liu (Beihang University) Traction (2026-06): 0 citations (arXiv)

Problem
Almost all vision-language manipulation research targets rigid robotic arms, whose fixed morphology limits adaptability in cluttered or confined spaces. Soft continuum arms are an appealing alternative because their deformability lets them squeeze around obstacles, but they bring hard new challenges: unreliable proprioception (the arm's true shape/state is hard to sense) and distributed, low-level continuum actuation rather than a few rigid joints. No benchmark previously existed to study vision-language manipulation under these conditions.
Method
ManiSoft contributes a tailored simulator plus benchmark. The simulator models the soft arm as a Cosserat rod driven by external torque, coupling realistic soft-body dynamics with contact-rich interaction via an elastic force constraint — a stability term (R_s, weighted by β) that the paper ablates to tune control stability.

ManiSoft defines four tasks of escalating difficulty:
- Collecting (COLL): gather a designated object into a container — basic trajectory control and end-effector coordination.
- Alignment (ALN): position a target to a specified 6-DoF pose — fine-grained orientation.
- Stacking (STK): assemble tableware largest-to-smallest into a stable pile — precise, elevated, contact-rich control.
- Arrangement (ARR): place objects in a specified spatial configuration — perception, spatial reasoning, and obstacle avoidance.
Data generation is automated: a high-level planner decomposes each task into waypoints, then a low-level reinforcement-learning policy emits torque commands to track them, producing expert trajectories at scale. The dataset contains 6,300 scene–trajectory pairs (2,100 clean + 4,200 randomized), with ~40 language instructions per scene, split 4:1 train/test.
Results
Three representative policies are benchmarked — Diffusion Policy (DP), RDT, and OpenVLA-OFT — in clean and randomized settings (success rate / steps):
- DP: 31.6% avg success (~520 steps); OpenVLA-OFT: 30.4% (~527 steps) — comparable, both ~400M params.
- RDT (~1B params): only 9.2% avg success despite ~496 steps, a large gap attributed to model capacity.
- All models show relatively promising clean-scene results but a substantial drop under randomization (e.g., in the ARR task OpenVLA-OFT falls from 31.3% clean to 13.7% randomized average).
Visualization analysis attributes failures to two root causes: inaccurate visual estimation of proprioceptive state and limited exploitation of the arm's deformability for adaptive obstacle avoidance (plus a "stop-moving" failure mode).
Significance
ManiSoft is the first benchmark to bring vision-language manipulation to soft continuum robots, exposing failure modes — proprioceptive ambiguity and under-used compliance — that rigid-arm benchmarks cannot surface. Its Cosserat-rod simulator and 6,300-scene RL-generated dataset give the community a testbed for policies that must reason about deformable morphology.
Links
- arXiv: 2605.18617
- ICML 2026: https://icml.cc/virtual/2026/poster/63416
← Back to ICML-2026