ICML 2026 SoMA - Heungwoo/research GitHub Wiki
SoMA — A 3D-Gaussian-Splat neural simulator for robot-conditioned soft-body manipulation
Venue: ICML 2026 (Poster) Category: World Model Traction (2026-06): 3 citations (arXiv)

Problem
Simulating deformable objects under rich interactions is a fundamental challenge for real-to-sim robot manipulation, because the dynamics are jointly driven by environmental effects and robot actions. Existing simulators rely on predefined physics or data-driven dynamics without robot-conditioned control, which limits accuracy, stability, and generalization. The two baselines SoMA targets are PhysTwin (differentiable simulator that relies on explicit point tracking and degrades under occlusion) and GausSim (state-based neural simulator that drifts over long horizons).
Method
SoMA is a 3D Gaussian Splat simulator that couples deformable dynamics, environmental forces, and robot joint actions in a single unified latent neural space, enabling end-to-end real-to-sim simulation. From synchronized multi-view RGB (three cameras) and the robot joint state R_t, the object is represented as Gaussian splats G_0 and rolled out by an action-conditioned simulator G_t = φ_θ(G_{t-1}, G_{t-2}, R_t) (Eq. 4), rendered under known camera poses and trained end-to-end by an image-reconstruction loss across time and views (Eq. 5). Three components make this work:
- Scene-to-simulation (R2S) mapping lifts heterogeneous reconstructions, robot kinematics, and physical reference frames into one unified simulation space, recovering a global scale factor by enforcing metric consistency so joint-space actions can directly drive object dynamics.
- Hierarchical graph-based simulator models robot–object–environment interactions through structured interaction graphs, enabling localized contact reasoning with global physical consistency.
- Multi-resolution training optimizes dynamics across temporal and spatial scales, paired with a blended supervision scheme combining occlusion-aware image losses and physics-inspired consistency constraints, for stable long-horizon rollouts.

Results
Headline claim: SoMA "improves resimulation accuracy and generalization on real-world robot manipulation by 20%." Data is collected on an ARX-Lift platform over four deformable objects (rope, doll, cloth, T-shirt), 640×480 RGB at 30 FPS, 30–40 sequences per object, 7:3 train/test split. On the main resimulation / generalization table, SoMA beats both baselines on every metric — Resimulation: Abs Rel 0.089, RMSE 0.124, PSNR 33.51, SSIM 0.971, LPIPS 0.055 (vs PhysTwin 0.102/0.150/28.77/0.947/0.086 and GausSim 0.115/0.155/31.69/0.945/0.092); Generalization: Abs Rel 0.112, RMSE 0.137, PSNR 32.89, SSIM 0.968, LPIPS 0.062. On the harder T-shirt folding task, SoMA reaches PSNR 27.57 / SSIM 0.896 / LPIPS 0.128 vs PhysTwin 22.85 / 0.842 / 0.198. An ablation on history length finds the best PSNR at k=10 (32.73).
Significance
SoMA pushes real-to-sim past rigid bodies and end-effector-trajectory conditioning: by binding deformable dynamics directly to robot joint actions inside a Gaussian-splat latent space, it produces controllable, occlusion-robust, long-horizon simulation of cloth and other soft bodies — a key ingredient for scalable manipulation world models and policy training without hand-built physics.
Links
- arXiv: 2602.02402
- ICML 2026: https://icml.cc/virtual/2026/poster/63706
← Back to ICML-2026