RSS 2026 Reading List - Heungwoo/research GitHub Wiki
RSS 2026 β VLA & Manipulation: Themed Reading List
The 8-theme annotated reading list for the RSS 2026 survey. Each theme opens with a π§ insight, then lists the in-scope papers with every entry linked to its per-paper page. (Split out of the survey so both pages render quickly.)
β Back to RSS 2026 survey Β· Full paper index Β· Home
1. RL & self-improvement for VLAs
π§ Insight β the improvement loop is now fully stocked at every layer: a policy-side interface (advantage tokens), reward sources (learned reward models, VLM verifiers, binary outcomes), and training infrastructure (unified RL frameworks). The differentiator has shifted from "can we RL a flow policy at all" (solved three ways) to "where does the reward signal come from" β RECAP bets on cheap binary advantage labels + human corrections, Robometer on trajectory comparisons, the verifier line on visual judgment. What no paper provides yet: a shared metric of improvement efficiency (Ξsuccess per robot-hour), so the approaches cannot be compared head-to-head.
- Ο*0.6 + RECAP (#87) β the flagship; advantage-conditioned RL over heterogeneous deployment experience; 2Γ throughput / ~Β½ failures on laundry, box assembly, espresso.
- RLux-VLA (#89) β unified, efficient RL-for-VLA framework (fair comparison across architectures/algorithms).
- Continual RL fine-tuning for long-lived VLAs (#86)
- Self-Improving Policy w/ Compositional World Model (#12)
- BCβQ-functions (#153)
- TMRL (#208)
- Visual verification generatorβverifier (#79)
- Robometer reward models (#140)
- Latent policy steering via one-step flow (#152).
2. Human data & cross-embodiment transfer
π§ Insight β three mutually incompatible mechanisms now coexist (let transfer emerge via co-training diversity Β· decouple representation-from-video / control-from-robot-data Β· synthesize robot data from video), and the right choice tracks two variables: embodiment distance (gripper fleets β emergence; 36-DoF humanoids β decoupling) and robot-data budget (small β decoupling; none β synthesis). All camps converge on the same substrate β hand-pose-annotated egocentric video β making that the no-regret investment regardless of camp. Full analysis: Review-Human-Video-Transfer.
- Emergence of Human-to-Robot Transfer (#72, Physical Intelligence) β transfer emerges with pre-training diversity; ~2Γ on human-only generalization settings.
- Ξ¨β (#21, in-depth) β the decoupling counter-thesis; open humanoid foundation model.
- HoMMI (#205, Stanford/TRI) β UMI + egocentric sensing β robot-free whole-body mobile manipulation.
- DexImit (#3)
- TactAlign (#6)
- EgoVerse (#92)
- EgoHumanoid (#204)
- SoftAct (#202)
- UMI-Underwater (#9)
- LAP language-action pretraining for zero-shot β¦ (#203)
- X-DiffVLA (#81).
3. Video-action & world models
π§ Insight β the unifying technical move is choosing the right prediction space: LDA-1B's structured DINO latents and mimic-video's video-model latents both avoid pixel-appearance modeling so dynamics becomes learnable from non-expert and even actionless data. The challengers win biggest exactly where BC data is scarcest (dexterous +48%), suggesting dynamics pre-training is a data-scarcity remedy more than a universal upgrade. The open flank: none of these models yet matches VLM-backbone instruction-following depth. Full analysis: Review-World-Models.
- mimic-video (#77) β the VAM-vs-VLA argument; 10Γ sample efficiency.
- LDA-1B (#210) β unified world model at 1B; +21/48/23% over Ο0.5 on contact-rich/dexterous/long-horizon.
- CauVA causal world modeling (#16)
- Act2Goal (#15)
- Interactive World Simulator (#18)
- HAIC dynamics-aware humanoid interaction (#13)
- Simulation Distillation (#17)
- long-horizon EWM consistency (#14)
- memory retrieval for visuomotor policies (#10).
4. Tactile / contact-grounded manipulation
π§ Insight β touch graduated from input modality to predicted state: the winning designs make the policy forecast future contact (ViTacFormer) or command contacts directly through a learned contact-to-controller map (CGP), effectively building a world model in contact space where pixel-space models are weakest. Hardware and simulation co-evolve in lockstep (deformation-independent sensing, shear-accurate sim-to-real) β a sign the bottleneck is shifting from sensor availability to representation design. Full analysis: Review-Tactile-VLA.
- Contact-Grounded Policy (#5)
- ViTacFormer (#129) β the two dexterous visuo-tactile flagships.
- TouchGuide (#78)
- Contact-Anchored Policies (#141)
- Force Policy (#128)
- TACTIC whole-arm (#60)
- Semantic Contact Fields (#4)
- HydroShear (#155)
- sensors: LightTact (#193), super-res multi-axis skin (#199), ECT proximity+grasp (#197).
5. Dexterous hands & cross-hand generalization
π§ Insight β the arm world's canonicalization playbook arrived at 20+ DoF: align the morphology first (anatomical graphs or canonical URDFs), and a single policy then transfers zero-shot at 80%+ to unseen hands β retargeting modules disappear entirely. Division of labor is stabilizing: sim-to-real RL keeps pushing the skill frontier (catching, in-hand reorientation, tool use) while IL + canonical representations carry task breadth; the two lines have not yet been combined in one system. Full analysis: Review-Cross-Embodiment Β· Review-Dexterous-Manipulation.
- DexGrasp-Zero (#122)
- One Hand to Rule Them All (#124) β the cross-embodiment hand pair.
- DexEvolve (#59)
- reactive catching sim-to-real (#148)
- ViserDex 3DGS in-hand reorientation (#150)
- SimToolReal (#151)
- extrinsic dexterity in clutter (#149)
- CRAFT hand (#192).
6. Humanoid loco-manipulation & whole-body control
π§ Insight β decomposition currently beats end-to-end: delegating balance to an RL System-0 and formulating manipulation targets in the world frame (HiWET) rather than the body frame is what makes long-horizon loco-manipulation trainable at all, and Ξ¨β shows the data argument is efficiency (right data, staged) rather than volume. The cost is capped agility β no VLA-driven stack yet exploits whole-body contact or dynamic bracing. Full analysis: Review-Humanoid-VLA.
- Ξ¨β (#21) β see in-depth review.
- HiWET world-frame EE tracking (#30)
- HAIC (#13)
- TeleGate teleop (#25)
- OmniXtreme high-dynamic tracking (#31)
- perceptive parkour (#20)
- HUSKY skateboarding (#19)
- X-Loco (#22)
- pixels-to-locomotion (#27)
- EgoHumanoid (#204)
- HoMMI (#205).
7. Action representation & inference mechanics
π§ Insight β two orthogonal levers emerged for the latency problem: make the token space ordered so decoding is anytime (OAT), or make chunk continuation native to training so smoothness needs no inference patch (Legato). Chunk-boundary handling is now a subfield with its own metrics (NSPARC smoothness), and the consistent lesson across papers is that inference-time fixes lose to trained-in properties. Full analysis: Review-Realtime-Execution.
- OAT (#75) β ordered action tokenization: compression + total decodability + causal order, with anytime prefix decoding.
- Legato (#58) β training-time native continuation for chunked flow policies; beats RTC by ~10% on smoothness and completion time.
- AR-VLA autoregressive action expert (#85)
- Action-to-Action flow matching (#209)
- SkillVLA skill reuse for dual-arm (#82)
- BagelVLA interleaved VLA generation (#83)
- GuidedVLA attention specialization (#84)
- StereoVLA (#88)
- PointACT (#73)
- steerable VLA hierarchies (#74)
- DISC instruction/state decoupling (#147)
- OrderedMoE / semantically structured MoE (#57)
- ENAP neural automata (#144)
- key-history-frame long-context IL (#201).
8. Evaluation, benchmarks & data infrastructure
π§ Insight β evaluation became a systems discipline with three pillars: reconstruct reality (PolaRiS's validated real-to-sim), decompose capability (LIBERO-X's graded pyramid, RoboTwin-IF's language probes), and get the statistics right (beyond binary success). The meta-shift: benchmark papers now must validate their own correlation with reality β an obligation that didn't exist in 2025. The LBM study anchors the data side: recipe questions are settled empirically, at 89-policy scale. Full analysis: Review-VLA-Evaluation.
- PolaRiS (#62)
- LIBERO-X (#97) β the two evaluation flagships.
- LBM co-training study (#7, TRI) β the 89-policy data-modality anchor study.
- MolmoSpaces (#91)
- RoboLab (#96)
- OopsieVerse (#98)
- Beyond Binary Success (#76)
- Betting for sim-to-real eval (#90)
- GS-Playground (#93)
- RoboVista (#95)
- R2RGen real-to-real 3D data generation (#121)
- One-shot bimanual demo synthesis (#1)
- PolaRiS-style offline eval (#154).
β Back to RSS 2026 survey Β· RSS-2026-Papers Β· Home