RSS 2026 X DiffVLA - Heungwoo/research GitHub Wiki
X-DiffVLA: X-Embodied Diffusion Action Heads for Vision-Language-Action Models
Venue: RSS 2026 (Sydney, Jul 13–17) · Session: VLA Models · paper #81 Authors: Boyu Li, Chaoyi Xu, Haoqi Yuan, Xinrun Xu, Börje F. Karlsson, Dongbin Zhao, Haoran Li, Zongqing Lu arXiv: 2605.25044 · program page
Summary compiled from the arXiv paper (v2); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Architectural overview: a VLA backbone plus proprioceptive soft prompts feed a unified diffusion action head whose action space is partitioned into base / gripper / dexterous-hand segments. Embodied Forcing (bottom left) injects morphology-dependent pyramid noise levels during denoising; Morphological Tree Diffusion (right) runs a Monte-Carlo-tree-style selection / expansion / simulation / backpropagation search over denoising paths, keeping high-value guided paths and rejecting others.
Problem
Pre-trained VLAs still need embodiment-specific fine-tuned action heads, which blocks knowledge transfer between robots that share a base but differ in end-effectors (parallel grippers vs dexterous hands). The paper targets a unified cross-embodied action head that lets heterogeneous demonstrations reinforce each other.
Method
X-DiffVLA post-trains a diffusion-based action head over a unified action space partitioned along a morphological tree ([base, gripper, hand], with zero-padding for absent segments). Two techniques make it work: Embodiment Forcing (EBF), a classifier-free-guidance-style noise-initialization scheme (pyramid noise, decay coefficient δ and window size Wn; best δ = 0.5, Wn = 4) that implicitly steers generation toward embodiment-specific functional components; and Morphological Tree Diffusion (MPTD), which adapts Monte Carlo Tree Search to the denoising process (selection, expansion, simulation with meta-actions, backpropagation) to strengthen behavioral correlations across end-effectors. Notably, EBF helps the diffusion head (+13.3 points) but not a flow-matching head, which the authors attribute to flow matching's simplified probability paths.
Results
On RoboCasa (GR00T-dataset benchmark, 30 task categories per end-effector; Panda, Robotiq-85, Inspire hand), X-DiffVLA averages 64.5% success vs π0.5 49.2%, X-VLA 47.6%, π0+FAST 43.1%, GR00T-N1 39.5% — the claimed +15.3% over the best baseline — while cutting both mobility and operation failure counts. On Isaac Gym dexterous grasping (Panda/Inspire/Shadow, DemoGrasp protocol) it averages 71.0% vs π0.5's 58.5% (+12.5%). Real-world FR3 pick-up tasks (GELLO + Manus glove teleop, 10 trajectories × 5 objects) give 63.5% average vs 57.0% for π0.5; ablations show removing EBF drops RoboCasa average to 43.5% and removing MPTD to 54.1%.
Significance
A concrete architectural answer to cross-embodiment transfer at the action-head level rather than the pre-training level, with an interesting negative result for flow-matching heads — relevant to Review-VLA-Architecture and Review-Cross-Embodiment.
← Back to RSS 2026 survey · RSS-2026-Papers · Home