ICLR 2026 RL Tokens - Heungwoo/research GitHub Wiki
Venue: Physical Intelligence research / tech report ยท Date: March 19, 2026 Authors: Charles Xu, Jost Tobias Springenberg, Michael Equi, Ali Amin, Adnan Esmail, Sergey Levine, Liyiming Ke Category: RL for VLA (production-flavored) Trend tag: Trend 3 (RL fine-tuning closes the gap) (Not an ICLR 2026 paper, but tightly relevant to the same conversation โ listed here because most ICLR 2026 RL work compares against the ฯ0.6 lineage.)
flowchart LR
Obs[Observation] --> P06[ฯ0.6<br/>FROZEN VLA]
P06 --> Act[Action chunk<br/>via flow-matching head]
P06 --> RLT[NEW: RL Token output<br/>compact summary of internal state]
RLT --> Actor[Tiny RL actor]
RLT --> Critic[Tiny RL critic]
Actor -.adjust.-> Act
Env[Real robot rollout] --> Reward
Reward --> Critic
Critic -.policy gradient.-> Actor
Actor -. minutes-scale online RL .-> Better[Improved precise behavior]
The RL Token acts as a compressed interface between the large frozen VLA and a lightweight RL policy. See the official PI page (link below) for the authors' own diagrams of the architecture and the four task videos.
Even with ฯ0.6 in production, the last millimeter of contact-rich precise tasks โ fully seating a screwdriver, threading a zip-tie, inserting an ethernet plug, plugging a power cord โ remains the failure mode that defines the gap between "demo" and "deployment." Standard recipes for closing this gap are unworkable in practice:
- Full fine-tuning of a 7B+ VLA on real-robot RL takes days and destabilizes the backbone.
- Imitation collection at the precision needed (sub-mm) is brutally expensive in human teleop hours.
- RECAP-style outcome conditioning (ฯ*0.6 + RECAP) works at task scope but is not optimized for fast contact-level adaptation.
Add a single special "RL token" output to the frozen ฯ0.6. The RL token is a compact vector that summarizes the high-dimensional internal embeddings of the VLA into something a small downstream RL policy can consume. Crucially, it is trained as an information bottleneck: an encoder compresses the VLA's final-layer embeddings into the single <rl> token, and an auxiliary decoder transformer must autoregressively reconstruct the original VLA embeddings from that token alone. This reconstruction objective forces the token to retain the task-relevant visual/contextual state, making it a sufficient (not lossy) summary for downstream RL. Then:
- A tiny actor consumes the RL token and produces a residual / steering signal that nudges the action chunk. The actor receives the VLA's predicted action chunk as input (it learns to edit rather than replace it) and is regularized toward that reference action, with reference-action dropout.
- A tiny critic estimates value from the same RL token.
- Online rollouts on the real robot update only the actor + critic (the VLA stays frozen).
Because only the lightweight head trains, online RL converges in a few hours rather than days, and the precision-relevant adjustments happen entirely at the contact-level interface without disturbing the backbone's general capabilities.
Across four challenging contact-rich manipulation tasks โ screwdriving, zip-tying, ethernet insertion, power-cord plugging โ RLT:
- Speeds up the most precise stages by up to 3ร
- Can surpass human teleoperation speed on those stages โ on Ethernet insertion, half the trials of the final RL policy are faster than any teleoperated demo (median 66 vs. 146 timesteps), trained on ~2 hours / 15 min of total robot data
- Achieves convergence in minutes (e.g., the M3-screw phase from as little as 15 minutes of real-world data) rather than days
- Leaves the underlying ฯ0.6 capabilities intact (no catastrophic forgetting)
RLT is the production-quality answer to the residual-RL story that ICLR 2026 papers like PLD and RFS tell at the benchmark level. The pattern is the same โ freeze the big model, train a tiny adapter with online RL โ but RLT is engineered for real robots doing real precise tasks, on the time budget of an actual deployment cycle (hours, not days).
It also defines the "compact RL interface" pattern as a first-class architectural feature: a single learned token whose entire job is to be the bridge between a frozen foundation policy and an online RL loop. Expect this to be copied widely in 2026โ2027 โ the engineering ergonomics are dramatically better than residual-policy approaches that need their own rich observation interface.
- Project page: https://www.pi.website/research/rlt
- Paper PDF: https://www.pi.website/download/rlt.pdf
- PI blog: https://www.pi.website/blog
- Independent explainer (Mochan Shrestha): https://mochan.org/posts/rlt/
- News writeup (Humanoids Daily): https://www.humanoidsdaily.com/news/the-last-millimeter-physical-intelligence-unveils-rl-tokens-for-hyper-fast-precision
- ฯ0.6 (the frozen base policy)
- ฯ*0.6 + RECAP (outcome-conditioned alternative for the same base)
- PLD (residual-RL ICLR-2026 counterpart)
- RFS (residual-RL for dexterity)
- SimpleVLA-RL (the scaling-focused RL recipe)
- RL for VLA โ section landing