RSS 2026 VLA Manipulation Survey - Heungwoo/research GitHub Wiki
RSS 2026 — VLA & Manipulation Survey
Venue: Robotics: Science and Systems (RSS) 2026 · Sydney, Australia · July 13–17, 2026 Compiled from the official accepted-papers program (roboticsconference.org/program/papers, July 2026 snapshot; all abstracts verified against the per-paper program pages).
At a glance
- 210 accepted papers across 21 oral sessions — RSS stays deliberately small and selective (contrast ICRA 2026's ~2,820).
- Manipulation is the center of gravity: ~116 of 210 papers (~55%) fall in scope for this wiki — three dedicated Manipulation sessions (27), a VLA Models session (9), three Imitation Learning sessions (28), Humanoids (13), World Models & Memory (9), plus manipulation-heavy papers throughout RL, Datasets & Benchmarks, Robot & Sensor Design, Perception, and Planning.
- The flagship: π*0.6 + RECAP — Physical Intelligence's "a VLA That Learns From Experience" — headlines the VLA session: hours-long laundry folding in real homes, factory box assembly, and espresso making, with RL from deployment experience more than doubling throughput and roughly halving failure rates on the hardest tasks.
- Session census (in-scope): Manipulation 1–3 · VLA Models · Imitation Learning 1–3 · Humanoids · World Models & Memory · RL (7/9 papers dexterity- or manipulation-relevant) · Datasets & Benchmarks (9/9 relevant) · plus hands/tactile hardware in Robot & Sensor Design.
The research flow — what RSS 2026 says the field is doing
Where ICLR 2026 was the architecture conference, CVPR 2026 the perception conference, and ICRA 2026 the deployment conference, RSS 2026 is the improvement-loop conference: the dominant question is no longer "how do we build a VLA?" but "how does a deployed policy get better — from its own experience, from human video, from touch, and how do we even measure that it did?"
Six threads carry the program:
- RL-from-experience reaches production VLAs. π*0.6/RECAP (#87) is the anchor: advantage-conditioned policies ingesting demonstrations, rollouts, and corrections at deployment scale. Around it, a full ecosystem forms — RLux-VLA (#89, a unified RL-for-VLA training framework), continual-learning VLAs via reinforcement fine-tuning (#86), self-improving policies through compositional world models (#12), Q-functions extracted from BC for on-robot RL (#153), diffusion-timestep-modulated pretraining for exploration (#208), generator–verifier autonomous improvement (#79), and Robometer's trajectory-comparison reward models (#140). The "VLA improvement loop" is now a first-class research object, not a workshop topic.
- Human video → robot skill transfer matures — and bifurcates. One camp shows emergence: PI's co-training study (#72) finds human-to-robot transfer emerges once robot pre-training is diverse enough, nearly doubling generalization on human-only settings. The other camp shows decoupling: Ψ₀ (#21) argues co-training human and humanoid action distributions is suboptimal and wins with staged training (800 h human video + 30 h robot data beating 10× corpora by >40 pp). Around them: DexImit (#3, monocular human video → bimanual dexterity), TactAlign (#6, tactile-glove H2R transfer), EgoVerse (#92, the global egocentric dataset — also in Qwen-RobotManip's corpus), EgoHumanoid (#204), HoMMI (#205, robot-free whole-body mobile-manipulation data), SoftAct (#202), and UMI-Underwater (#9). The shared bet: the demonstration bottleneck breaks on humans, not teleop fleets.
- Video/world models challenge the VLA backbone itself. mimic-video (#77) makes the frontal argument — VLM backbones are "blind to physical causality"; a video-pretrained backbone + inverse-dynamics action decoder yields 10× sample efficiency. LDA-1B (#210) scales a unified world-model formulation (dynamics + policy + forecasting in DINO latent space) to 1B on 30k hours and beats π0.5 by up to 48% on dexterous tasks. CauVA (#16), Act2Goal (#15), the Interactive World Simulator (#18), and simulation-distilled world models (#17) round out a session arguing that dynamics knowledge, not just semantics, belongs in pre-training — the same thesis as Review-World-Models's WAM line.
- Touch and contact become policy-level citizens. Not sensors bolted onto VLAs (the ICRA framing) but contact as representation: Contact-Grounded Policy (#5) predicts coupled robot-state + tactile trajectories and grounds them through a contact-consistency map; ViTacFormer (#129) learns cross-attended visuo-tactile latents with autoregressive tactile prediction (~50% higher success, 11-stage long-horizon dexterity); TouchGuide (#78) steers pretrained visuomotor policies with touch at inference time; Contact-Anchored Policies (#141) replace language conditioning with contact points; Force Policy (#128), TACTIC (#60, whole-arm contact), Semantic Contact Fields (#4), and HydroShear (#155, shear-accurate tactile sim-to-real) complete the picture.
- Dexterous hands go cross-embodiment. The hand community adopts the cross-embodiment agenda wholesale: DexGrasp-Zero (#122) hits 85% zero-shot grasp success on unseen hands via morphology-aligned graphs; One Hand to Rule Them All (#124) canonicalizes hand URDFs into a parameterized representation with a learnable morphology latent (81.9% zero-shot on an unseen 3-finger hand); DexEvolve (#59) scales grasp-synthesis data evolutionarily; sim-to-real RL delivers reactive catching (#148), in-hand reorientation from monocular RGB via 3DGS (#150), zero-shot tool use (#151), and extrinsic dexterity in clutter (#149). Hardware keeps pace (CRAFT tendon hand #192, LightTact #193, super-resolution skin #199).
- The evaluation crisis gets institutional answers. The community's response to benchmark distrust (cf. RobotManip's manifesto): PolaRiS (#62) turns phone scans into simulation evals that actually rank real-world generalist policies (600 real + 93k sim rollouts of validation); LIBERO-X (#97) rebuilds LIBERO with hierarchical perturbation protocols; RoboLab (#96), MolmoSpaces (#91, AllenAI's open manipulation+navigation ecosystem), OopsieVerse (#98, damage-aware safety), Beyond Binary Success (#76, statistically rigorous policy comparison), Betting-based sim-to-real evaluation (#90), and offline policy evaluation via discounted liveness (#154). Evaluation methodology is now a publishable contribution class of its own.
Cross-cutting observation: the TRI LBM co-training study (#7 — 89 policies, 58k sim + 2,835 real rollouts) lands as the definitive empirical anchor for the field's data-mixing questions, confirming at scale what the ICLR study began: VL and cross-embodiment co-training deliver cumulative generalization gains, while discrete action tokens again show no significant benefit — the third independent nail (after LBM-1 and the Qwen program's omission) in the discrete-action-token coffin.
Themed reading list
→ RSS 2026 — Themed Reading List — 8 themes (RL & self-improvement · human data & cross-embodiment · video/world models · tactile · dexterous hands · humanoid loco-manipulation · action representation · evaluation & benchmarks), each with a 🧠 insight and every in-scope paper linked. (Moved to a dedicated page so this survey renders quickly.)
Full in-scope session tables
→ RSS 2026 — Full In-Scope Paper Index — all 116 papers by session (VLA Models · Manipulation 1–3 · Imitation Learning 1–3 · Humanoids · World Models & Memory · RL · Datasets & Benchmarks · hands/tactile picks), every entry linked to its per-paper page. (Moved to a dedicated page so this survey renders quickly.)
How RSS 2026 relates to the other 2026 venues
- vs ICRA 2026: ICRA carried the sensor/deployment breadth (728 in-scope papers); RSS concentrates the method frontier — nearly every RSS manipulation paper would headline an ICRA session.
- vs ICLR/ICML 2026: the ML venues debated architectures and representations; RSS closes the loop with hardware evidence — RECAP, Ψ₀, ViTacFormer, and the LBM study are all real-robot-first results.
- Continuity: RTC (NeurIPS 2025) spawns Legato and Ψ₀'s training-time variant; LBM co-training (ICLR) gets its large-scale sequel (#7); EgoDex becomes the pre-training corpus of choice (Ψ₀); the Qwen-RobotManip evaluation manifesto finds its infrastructure answers (PolaRiS, LIBERO-X, RoboLab); and the discrete-action-token negative result is now triply replicated.
Related pages
- Per-paper pages: PI-RECAP
- Review-Psi0
- RSS-2026-LDA-1B
- RSS-2026-Human2Robot-Emergence
- RSS-2026-mimic-video
- RSS-2026-LBM-Cotraining-Study
- RSS-2026-ViTacFormer
- RSS-2026-DexGrasp-Zero
- RSS-2026-One-Hand
- RSS-2026-Contact-Grounded-Policy
- RSS-2026-PolaRiS
- RSS-2026-LIBERO-X
- RSS-2026-OAT
- RSS-2026-Legato
- RSS-2026-HoMMI
- Venue hub: RSS
- Other venues: ICRA-2026-VLA-Manipulation-Survey
- ICLR-2026-VLA-Manipulation-Survey
- CVPR-2026-VLA-Manipulation-Survey
- ICML-2026
- CoRL-2025-VLA-Manipulation-Survey
- Topic reviews: Review-Dexterous-Manipulation
- Review-Tactile-VLA
- Review-Humanoid-VLA
- Review-Cross-Embodiment
- Review-World-Models
- RL
← Back to Home