ICML 2026 Seeing Realism from Simulation - Heungwoo/research GitHub Wiki

Seeing Realism from Simulation — Efficient video transfer for sim-to-real VLA data augmentation

Venue: ICML 2026 (Poster) Category: VLA Architecture / Diffusion-Flow Policy Affiliations: Chenyu Hui, Xiaodi Huang, Siyu Xu, Yunke Wang, Shan You, Fei Wang, Tao Huang, Chang Xu Traction (2026-06): 1 citation (arXiv)

Overall framework: coreset sampling selects important/diverse simulation trajectories, which are converted to realistic video strips via conditional video transfer and used for VLA training (Figure 1 from Hui et al., 2026)

Problem

VLA models are typically trained on large-scale real-world robot videos, which are costly and slow to collect. Simulated data is cheap and massively parallelizable, but suffers from a substantial visual domain gap and limited environmental diversity, so simulation-trained policies generalize poorly. The authors note this brittleness is well documented: on LIBERO-Plus, success rates drop from ~95% to under 30% under minor perturbations in object layout, camera angle, or lighting, and LIBERO-PRO reports near-zero accuracy when object positions and instructions change. This reveals that models memorize fixed action sequences rather than learning robust, semantic instruction-to-action mappings. The goal is to inject visual and environmental diversity into training data while strictly preserving task semantics and action trajectories — and to do so efficiently enough to scale.

Method

The pipeline transforms source (simulation) videos into visually diverse but semantically faithful videos in four stages: caption generation, caption rewriting, structured-condition extraction, and conditional video synthesis. Descriptive captions are extracted from source videos with a temporal video captioning model (VideoChat2), then rewritten by an LLM (Qwen3-8B) to diversify environments (backgrounds, lighting). Structured conditions are extracted via video semantic segmentation, and a conditional video transfer model (built on Cosmos-Transfer) synthesizes realistic videos conditioned on these cues.

Two efficiency mechanisms make augmentation practical at scale:

  • Velocity caching (diffusion feature-reuse): A segmented (three-stage) caching strategy reuses velocity predictions across adjacent denoising timesteps. The authors observe a stable phase where adjacent velocity predictions change minimally, enabling caching and reuse. This cuts generation time by over 60% with minimal accuracy loss.
  • Coreset sampling: A trajectory-level coreset formulation builds a graph that balances difficulty (measured by RDT-1B policy loss) and visual diversity (measured via Cosmos-Embed1 representations), selecting a compact, non-redundant subset (e.g., 50% of samples) to augment under limited compute.

Augmented data is combined with originals via either a mixture strategy (augmented + all originals) or a replacement strategy (selected coreset replaced by augmented versions).

Results

Experiments span RoboTwin 2.0, LIBERO, LIBERO-Plus, and a real AgileX robot, using RDT-1B, ACT, π₀, and π₀.₅.

  • RoboTwin 2.0: Improves RDT-1B by 8% (single- and multi-task, Easy/Hard settings using 50 demos/task with domain-randomized Hard mode covering clutter, lighting, textures, height variation).
  • LIBERO-Plus (spatial suite, Table 3): Boosts π₀ by 5.1% on this more challenging benchmark. Per-perturbation gains include light conditions for π₀ rising 75.0 → 78.7 (+3.7) and for π₀.₅ rising 94.5 → 97.9 (+3.2).
  • Real-world: On a physical AgileX Piper, evaluated on Stack Tape and Slot Pen across In-Distribution, Position Shift (OOD), and Background Shift (OOD), the augmentation yields consistent, substantial improvements under distribution shift.
  • Ablation: Velocity-cache vs. no-cache variants achieve near-identical downstream accuracy (both far above the original baseline), confirming that the >60% speedup does not compromise augmentation quality. Against RoboTransfer, the method achieves better geometric consistency (RMSE/Abs.Rel/Sq.Rel) and semantic alignment (VideoCLIP-XL).

Significance

The work shows that diffusion-based video transfer can convert cheap simulation rollouts into diverse, realistic training data that meaningfully improves VLA robustness to real-world perturbations, while the velocity-caching and coreset contributions make such augmentation affordable at dataset scale — a practical lever for closing the sim-to-real gap without collecting more real data.

Links

← Back to ICML-2026