ICML 2026 ManiSoft - Heungwoo/research GitHub Wiki

ManiSoft — A vision-language manipulation benchmark for soft continuum robots

Venue: ICML 2026 (Poster) Category: Benchmark Affiliations: Ziyu Wei, Luting Wang, Chen Gao, Li Wen, Si Liu (Beihang University) Traction (2026-06): 0 citations (arXiv)

ManiSoft overview: a vision-language manipulation benchmark built on soft continuum arms with four deformable-control tasks (Figure 1 from Wei et al., 2026)

Problem

Almost all vision-language manipulation research targets rigid robotic arms, whose fixed morphology limits adaptability in cluttered or confined spaces. Soft continuum arms are an appealing alternative because their deformability lets them squeeze around obstacles, but they bring hard new challenges: unreliable proprioception (the arm's true shape/state is hard to sense) and distributed, low-level continuum actuation rather than a few rigid joints. No benchmark previously existed to study vision-language manipulation under these conditions.

Method

ManiSoft contributes a tailored simulator plus benchmark. The simulator models the soft arm as a Cosserat rod driven by external torque, coupling realistic soft-body dynamics with contact-rich interaction via an elastic force constraint — a stability term (R_s, weighted by β) that the paper ablates to tune control stability.

Soft-arm modeling in the ManiSoft simulator: the soft body is a Cosserat rod moving under external torque with contact handling (Figure 2 from Wei et al., 2026)

ManiSoft defines four tasks of escalating difficulty:

  • Collecting (COLL): gather a designated object into a container — basic trajectory control and end-effector coordination.
  • Alignment (ALN): position a target to a specified 6-DoF pose — fine-grained orientation.
  • Stacking (STK): assemble tableware largest-to-smallest into a stable pile — precise, elevated, contact-rich control.
  • Arrangement (ARR): place objects in a specified spatial configuration — perception, spatial reasoning, and obstacle avoidance.

Data generation is automated: a high-level planner decomposes each task into waypoints, then a low-level reinforcement-learning policy emits torque commands to track them, producing expert trajectories at scale. The dataset contains 6,300 scene–trajectory pairs (2,100 clean + 4,200 randomized), with ~40 language instructions per scene, split 4:1 train/test.

Results

Three representative policies are benchmarked — Diffusion Policy (DP), RDT, and OpenVLA-OFT — in clean and randomized settings (success rate / steps):

  • DP: 31.6% avg success (~520 steps); OpenVLA-OFT: 30.4% (~527 steps) — comparable, both ~400M params.
  • RDT (~1B params): only 9.2% avg success despite ~496 steps, a large gap attributed to model capacity.
  • All models show relatively promising clean-scene results but a substantial drop under randomization (e.g., in the ARR task OpenVLA-OFT falls from 31.3% clean to 13.7% randomized average).

Visualization analysis attributes failures to two root causes: inaccurate visual estimation of proprioceptive state and limited exploitation of the arm's deformability for adaptive obstacle avoidance (plus a "stop-moving" failure mode).

Significance

ManiSoft is the first benchmark to bring vision-language manipulation to soft continuum robots, exposing failure modes — proprioceptive ambiguity and under-used compliance — that rigid-arm benchmarks cannot surface. Its Cosserat-rod simulator and 6,300-scene RL-generated dataset give the community a testbed for policies that must reason about deformable morphology.

Links

← Back to ICML-2026