RSS 2026 Reading List - Heungwoo/research GitHub Wiki

RSS 2026 β€” VLA & Manipulation: Themed Reading List

The 8-theme annotated reading list for the RSS 2026 survey. Each theme opens with a 🧠 insight, then lists the in-scope papers with every entry linked to its per-paper page. (Split out of the survey so both pages render quickly.)

← Back to RSS 2026 survey Β· Full paper index Β· Home


1. RL & self-improvement for VLAs

🧠 Insight β€” the improvement loop is now fully stocked at every layer: a policy-side interface (advantage tokens), reward sources (learned reward models, VLM verifiers, binary outcomes), and training infrastructure (unified RL frameworks). The differentiator has shifted from "can we RL a flow policy at all" (solved three ways) to "where does the reward signal come from" β€” RECAP bets on cheap binary advantage labels + human corrections, Robometer on trajectory comparisons, the verifier line on visual judgment. What no paper provides yet: a shared metric of improvement efficiency (Ξ”success per robot-hour), so the approaches cannot be compared head-to-head.

2. Human data & cross-embodiment transfer

🧠 Insight β€” three mutually incompatible mechanisms now coexist (let transfer emerge via co-training diversity Β· decouple representation-from-video / control-from-robot-data Β· synthesize robot data from video), and the right choice tracks two variables: embodiment distance (gripper fleets β†’ emergence; 36-DoF humanoids β†’ decoupling) and robot-data budget (small β†’ decoupling; none β†’ synthesis). All camps converge on the same substrate β€” hand-pose-annotated egocentric video β€” making that the no-regret investment regardless of camp. Full analysis: Review-Human-Video-Transfer.

3. Video-action & world models

🧠 Insight β€” the unifying technical move is choosing the right prediction space: LDA-1B's structured DINO latents and mimic-video's video-model latents both avoid pixel-appearance modeling so dynamics becomes learnable from non-expert and even actionless data. The challengers win biggest exactly where BC data is scarcest (dexterous +48%), suggesting dynamics pre-training is a data-scarcity remedy more than a universal upgrade. The open flank: none of these models yet matches VLM-backbone instruction-following depth. Full analysis: Review-World-Models.

4. Tactile / contact-grounded manipulation

🧠 Insight β€” touch graduated from input modality to predicted state: the winning designs make the policy forecast future contact (ViTacFormer) or command contacts directly through a learned contact-to-controller map (CGP), effectively building a world model in contact space where pixel-space models are weakest. Hardware and simulation co-evolve in lockstep (deformation-independent sensing, shear-accurate sim-to-real) β€” a sign the bottleneck is shifting from sensor availability to representation design. Full analysis: Review-Tactile-VLA.

5. Dexterous hands & cross-hand generalization

🧠 Insight β€” the arm world's canonicalization playbook arrived at 20+ DoF: align the morphology first (anatomical graphs or canonical URDFs), and a single policy then transfers zero-shot at 80%+ to unseen hands β€” retargeting modules disappear entirely. Division of labor is stabilizing: sim-to-real RL keeps pushing the skill frontier (catching, in-hand reorientation, tool use) while IL + canonical representations carry task breadth; the two lines have not yet been combined in one system. Full analysis: Review-Cross-Embodiment Β· Review-Dexterous-Manipulation.

6. Humanoid loco-manipulation & whole-body control

🧠 Insight β€” decomposition currently beats end-to-end: delegating balance to an RL System-0 and formulating manipulation targets in the world frame (HiWET) rather than the body frame is what makes long-horizon loco-manipulation trainable at all, and Ξ¨β‚€ shows the data argument is efficiency (right data, staged) rather than volume. The cost is capped agility β€” no VLA-driven stack yet exploits whole-body contact or dynamic bracing. Full analysis: Review-Humanoid-VLA.

7. Action representation & inference mechanics

🧠 Insight β€” two orthogonal levers emerged for the latency problem: make the token space ordered so decoding is anytime (OAT), or make chunk continuation native to training so smoothness needs no inference patch (Legato). Chunk-boundary handling is now a subfield with its own metrics (NSPARC smoothness), and the consistent lesson across papers is that inference-time fixes lose to trained-in properties. Full analysis: Review-Realtime-Execution.

8. Evaluation, benchmarks & data infrastructure

🧠 Insight β€” evaluation became a systems discipline with three pillars: reconstruct reality (PolaRiS's validated real-to-sim), decompose capability (LIBERO-X's graded pyramid, RoboTwin-IF's language probes), and get the statistics right (beyond binary success). The meta-shift: benchmark papers now must validate their own correlation with reality β€” an obligation that didn't exist in 2025. The LBM study anchors the data side: recipe questions are settled empirically, at 89-policy scale. Full analysis: Review-VLA-Evaluation.


← Back to RSS 2026 survey Β· RSS-2026-Papers Β· Home