RSS 2026 Towards Long Lived Robots - Heungwoo/research GitHub Wiki

Towards Long-Lived Robots: Continual Learning VLA Models via Reinforcement Fine-Tuning

Venue: RSS 2026 (Sydney, Jul 13–17) · Session: VLA Models · paper #86 Authors: Yuan Liu, Haoran Li, Shuai Tian, Yuxing Qin, Yuhui Chen, Yupeng Zheng, Yongzhen Huang, Dongbin Zhao arXiv: 2602.10503 · program page

Summary compiled from the arXiv paper (v2); all numbers quoted from the paper. Trend context: RSS 2026 survey.

LifeLong-RFT overview (Figure 1 of arXiv 2602.10503, © the authors)

Figure 1 contrasts SFT post-training (numerous demonstrations, catastrophic forgetting, poor transfer) with LifeLong-RFT (limited demonstrations, on-policy chunk-level RL with a multi-dimensional process reward) in both multi-task and continual-learning regimes; the performance panel shows LIBERO continual-learning AUC gains (+15.1 to +35.9) and SimplerEnv multi-task gains, plus an SFT-forgetting vs RFT-preservation rollout comparison.

Problem

SFT is the default post-training route for VLAs but needs substantial task-specific data and causes catastrophic forgetting, blocking the evolution of VLAs into long-lived agents that keep acquiring skills. Existing RFT alternatives need environment-provided rewards (simulation, privileged state) or learned reward models prone to reward hacking — both requiring costly environment interaction.

Method

LifeLong-RFT (CASIA et al.) performs chunking-level on-policy RL with GRPO — group size 8, KL regularization to a reference policy — where each sampled action chunk is scored offline against expert demonstrations, so no environment interaction or pretrained reward model is needed. The multi-dimensional process reward r = ω·QACR + (1−ω)·CTAR + λ·FCR (ω=0.7, λ=0.1) combines: QACR, position-wise token matching in the FAST+ quantized action space; CTAR, an exponentially decaying L1 pose reward (α=5) plus a binary gripper reward (β=0.8) on decoded continuous chunks; and FCR, a binary format-validity reward. The base model is NORA-Long (discrete-action VLA with FAST+ tokenizer), fully fine-tuned at lr 1e-6 on 8 H20 GPUs.

Results

Multi-task: +3.5% average on SimplerEnv WidowX (65.5→69.0) and +4.4% on Google Robot (74.7→79.1) over the NORA-Long SFT baseline; on LIBERO the average reaches 95.6% (vs 91.8% SFT), the best among compared discrete- and continuous-action models; real-world Franka tasks improve +8.7% on average (+15% on Hang Chinese Knot). Continual learning (LOTUS-style base + lifelong stages, 10 demos per new task, 5-demo experience replay): on LIBERO the paper reports a 22% average success-rate gain over SFT, with AUC gains up to +35.9 on LIBERO-Goal and NBT cut to 1.5–12.8; real-world continual learning improves FWT +23.7 and AUC +31.7. New tasks are learned with only 20% of the data (e.g., 100% on "Pick Orange Juice" with 5 demos vs SFT's 98% with 50). Ablation: removing CTAR collapses performance (−90.9%); QACR and FCR contribute −2.8% and −2.6% respectively.

Significance

Transfers the LLM-community finding that on-policy RL resists forgetting into VLA post-training, with a fully offline, verifiable reward — a practical continual-learning recipe that needs neither simulators nor reward models. Currently limited to discrete-action VLAs. Related wiki threads: Review-VLA-Architecture · RL.

← Back to RSS 2026 survey · RSS-2026-Papers · Home