ICLR 2026 Sim2Real VLA - Heungwoo/research GitHub Wiki

Sim2Real-VLA β€” Zero-Shot Transfer of Synthesized Skills

Venue: ICLR 2026 Authors: Runyi Zhao, Sheng Xu, Ruixing Jin, Yueci Deng, Yunxin Tai, Kui Jia, Guiliang Liu (CUHK Shenzhen; DexForce; Shenzhen Loop Area Institute) Code: https://github.com/DexForce/EmbodiChain Category: VLA Architecture / Sim2Real Trend tag: Synthesized data + affordance structure for zero-shot transfer

Approach diagram

flowchart LR
  Real[Real teleop / human video] --> R2S[Real2Sim digital cousins<br/>+ object pose recovery]
  R2S --> Sim[Simulator scaling<br/>DR features ranked by GPT-5]
  Sim --> Skill[Auto skill acquisition<br/>VLM atomic-task decomposition<br/>+ generalized IK]
  Skill --> Train[Sim-only training]
  Train --> Plan[Planning: object-masked obs.<br/>β†’ Chain-of-Affordance q_0..q_K<br/>regressive transformer]
  Plan --> Act[Acting: arm-decoupled Ο€^l, Ο€^r<br/>DiT diffusion expert + FAST tokens<br/>+ EOS validation model]
  Act --> Robot[Agilex CobotMagic<br/>zero-shot real-world]
Loading

Problem

Recent work on synthesized robot data (X-Sim, DreamGen, RoboCasa, etc.) is cheap to scale but trains policies that overfit task-irrelevant simulator features (lighting, textures) and miss motion-critical dynamics, producing large Sim2Real gaps even when fine-tuned on the same synthesized data. Sim2Real-VLA argues this is a model-architecture problem, not a fidelity problem: redesigning the VLA so it can only "see" affordance-relevant features lets sim-trained skills transfer zero-shot to a real Agilex CobotMagic with no real-world fine-tuning gradient updates.

Detailed Method

The model is a planner-actor dual system tied together by a chain of affordances q = [q_0, ..., q_K], each affordance a set of geometrically structured keypoints corresponding to an end-effector pose for an atomic sub-task.

1. Planning system: Chain-of-Affordance.

  • Object-oriented observation: observation Γ΄_t = f_ΞΎ(e_0, e_1, ...) is rendered in simulation; a CNN-based mask predictor p_ΞΈ^R(m_t^i | o_t^ΞΎ, ..., o_{t-H}^ΞΎ) is jointly trained with the policy (Eq. 1). DR is applied via p^d over scene-level features (lighting, table texture, background, distractors, object location/orientation/texture/shape) and robot-level features (camera pose/orientation/FOV, initial EEF pose).
  • Strategic DR feature selection: a foundation model (GPT-5) ranks DR features and defines sampling ranges given task description + observation + sim config.
  • Flow-of-DR: at each time step, action-invariant features (lighting, textures, backgrounds) are resampled β€” unlike fixed-trajectory DR used by RoboCasa.
  • Affordance reasoning: a regressive transformer learns p_Ο•^A(q_{k,t}, ..., q_{K,t} | mΜ‚_t, o_t^ΞΎ, ..., l) (Eq. 2), predicting subsequent affordances from masked observations.

2. Acting system: Predictive control as affordance execution.

  • Arm-decoupled estimation: two independent controllers Ο€_Ο‰^l, Ο€_Ο‰^r, each conditioned only on its own wrist + top-down view and the arm-specific affordance target. Prevents cross-arm attention bleeding.
  • Tokenized action space: continuous actions β†’ DCT frequency-domain representation β†’ quantization β†’ BPE β†’ compact token sequence a^DCT (FAST tokenizer, Pertsch et al. 2025). Tokens are decoded by a DiT-style diffusion action expert.
  • Affordance validation model: a regressive transformer classifier takes masked obs + state + current target affordance and outputs whether the target is achieved β†’ decides next-affordance progression or repeat. Provides reactive failure recovery.

3. Automatic data generation (EmbodiChain pipeline).

  • Real2Sim (Sec. 4.3-1): per-object detection + retrieval to "digital cousins" from a sim asset library; articulated objects post-processed via CAD alignment; VLM-driven scene correction prompts (Listings 1–2) detect occlusion and revise the layout. Action trajectories retargeted from teleoperation or egocentric human video.
  • Generative scene scaling: sample DR features within Real2Sim priors.
  • Automatic skill acquisition: VLM decomposes task into atomic units, generates candidate grasp/manipulation poses, generalized IK fills in joint angles, terminal poses of each atomic task become affordance supervision q_k.

Architecture & training. DINOv2 visual encoder + T5-XXL language encoder. Action expert ~200 M params, 8-layer transformer, 256 hidden dim, 8 attention heads; two additional transformer blocks for affordance inference; multiple MLP adapters. Cosine-annealed LR (max 1e-5), 40k epochs, batch 8, ~36 GPU-hours, EMA for stability.

Hardware. Agilex CobotMagic bimanual robot. Calibrated CCTag-based binocular setup (3.8 mm error); wrist cameras from CobotMagic URDF.

Comprehensive Results

Main long-horizon table (Table 3 of paper). Six tasks, sim 50 runs / real 20 runs each, with per-task max-step caps:

Task (cap) Ο€0 (FwS) Sim/Real Ο€0-FAST Sim/Real GR00T N1.5 Sim/Real Sim2Real-VLA Sim/Real
Single-Arm Water Pour (200) 38/50, 6/20 31/50, 11/20 29/50, 9/20 46/50, 17/20
Dual-Arm Water Pour (250) 25/50, 5/20 30/50, 8/20 22/50, 7/20 47/50, 16/20
Table Rearrangement (250) 11/50, 4/20 23/50, 7/20 16/50, 4/20 44/50, 16/20
Items Hand-Over (400) 12/50, 4/20 10/50, 1/20 18/50, 3/20 31/50, 8/20
Basket Pick&Place (400) 15/50, 2/20 13/50, 3/20 9/50, 2/20 29/50, 9/20
Pan Open&Place (550) 12/50, 1/20 11/50, 3/20 17/50, 1/20 30/50, 7/20

Average real-world SR: 60.8% vs. best baseline ~25%, an absolute improvement of >35 pts. ACT and Diffusion Policy completely fail on the four hardest tasks (0/20).

Domain-shift generalization (Table 4). Background / object / table-texture variations Γ— 6 tasks Γ— 20 trials each. Single-Arm Pour: 17β†’17 (background), 17β†’16 (object), 17β†’17 (table), 17β†’16 (combined). Dual-Arm Pour and Table Rearrangement are similarly stable. The paper notes a counterintuitive observation: some tasks improve under domain shift because the shift removes spurious correlations.

Attention map analysis (Fig. 5). Without affordance guidance, RDT attention is broadly distributed across irrelevant background and unrelated robot joints. Sim2Real-VLA concentrates attention on motion-critical pixels for the current sub-task.

Ablation Studies

Arm-decoupled vs. joint learning (Table 8).

Task Joint Sim/Real/Steps Decoupled Sim/Real/Steps
Single-Arm Pour 0.86 / 0.75 / 178.6 0.92 / 0.85 / 174.6
Items Hand-Over 0.32 / 0.15 / 390.0 0.62 / 0.40 / 370.2

The effect is dramatic on the bimanual task: decoupling more than doubles real success rate. Authors attribute this to elimination of cross-arm visual interference.

Few-shot real-data adaptation (Table 9, Fig. 8). Sim2Real-VLA reaches 0.85 on Rearrangement and 0.50 on Basket with zero real demos (Sim Only) β€” beating Ο€0 / Ο€0-FAST that use 10 real eps. With 10 real eps via Sim-then-Real, Sim2Real-VLA hits 0.90 / 0.60. Notable "dip" at 5 eps (drops to 0.60) β€” authors interpret as the small real set perturbing the strong sim prior.

Training efficiency (Fig. 9). Converges in ~4 hours wall-clock; Ο€0 needs >10 hours. FLOPs estimated via FlashVLA formula; both metrics confirm Sim2Real-VLA is more efficient.

Segmentation Sim2Real transfer (Tables 10–11). Mean IoU vs. SAM masks on real images: 0.69–0.82 across tasks (lowest on Pan Open and Place). Even when transferred to an out-of-distribution camera setup (different placement, single-arm pouring), IoU vs. SAM = 0.78 β€” the segmentation module generalizes purely from DR.

Limitations stated by authors

The paper does not include a dedicated Limitations section. The Conclusion implicitly identifies extensions:

  • Tabletop-only; future work targets multi-agent collaboration and interactive environments beyond tabletop.
  • No RL refinement; the authors note RL integration on top of supervised affordance learning as a future direction.
  • The Real2Sim projection is described as "fully automatic" but depends on a curated digital-cousin asset library and a VLM corrective loop β€” degraded asset coverage will degrade scene fidelity.
  • The "zero-shot" framing is qualified in Sec. A.3: Real2Sim consumes a small set of real teleop/video trajectories to configure simulations but not as policy supervision β€” readers should interpret "zero-shot" accordingly.

Significance & Positioning

Sim2Real-VLA stakes out the model-redesign side of the Sim2Real debate. While Real2Sim photorealism work (NeRF-based digital twins, Cosmos) chases pixel-perfect simulation, this paper argues that affordance-driven attention + mask-conditioned observations + arm-decoupled execution are sufficient to filter task-irrelevant variance, so even a vanilla simulator's data trains a transferable policy.

Compared to:

  • X-Sim β€” relies on policy training inside high-fidelity sim; Sim2Real-VLA explicitly de-emphasizes fidelity.
  • DreamGen β€” generative video data; Sim2Real-VLA uses simulator with structural priors instead.
  • Genie Envisioner β€” world-model-driven; complementary.
  • OneTwoVLA β€” dual-system; Sim2Real-VLA's split is affordance-mediated.
  • VLBiMan, HWC-Loco β€” sibling CUHK Shenzhen / DexForce papers from the same group.

The >35 pt real-world SR margin under zero-shot deployment is the strongest empirical evidence to date for the "redesign the policy, not the simulator" position.

Links

Related pages

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️