CoRL 2025 TA VLA - Heungwoo/research GitHub Wiki
Venue: CoRL 2025 Β· arXiv: 2509.07962 Category: VLA Architecture Trend tag: Force as first-class policy signal
flowchart LR
V[Vision] --> VLM[VLM]
L[Language] --> VLM
VLM --> DEC[Action decoder]
T[Joint torque history] --> TOK[Summarize as<br/>SINGLE token]
TOK --> DEC
DEC --> A[Actions]
DEC --> TP[Predicted torque<br/>aux objective]
Most VLAs are vision-and-language only. Contact-rich tasks (wiping, insertion, pushing doors) quietly fail because the model can't tell success from failure until it sees a visual cue β far too late. Joint torque is the natural early signal, but naive injection (e.g., per-timestep torque tokens) destabilizes training and regresses non-contact tasks.
Systematic study of where and how to inject torque into a pretrained VLA (base model: Οβ, a flow-matching VLA). Two key findings:
- Where: place torque in the decoder as a single history-summary token (not per-timestep, not in the encoder) β this "preserves the original input pattern of the decoder" and balances informativeness with architectural stability.
- How (auxiliary objective): also predict torque as an auxiliary output via a joint actionβtorque diffusion loss (β_joint = β_action + Ξ²Β·β_torque), inspired by joint prediction-and-planning in autonomous driving. This pushes the model to build a physically grounded internal representation and further improves performance.
On 5 contact-rich tasks (20 trials each, Οβ baseline β TA-VLA): Button Pushing 5/20β18/20, Charger Plugging 0/20β17/20, USB Plugging 0/20β17/20, Socket Unplugging 16/20β19/20, Door Handle Turning 2/20β15/20. Raises the typically near-zero contact-rich success rates to ~75β95% with essentially no regression on the 5 regular (non-contact) tasks. Baselines include ACT, RDT-1B, and Οβ.
Part of CoRL 2025's force-as-first-class trend alongside UniFP (Best Paper), DexSkin, and Tactile Beyond Pixels. Delivers a clean, composable recipe that any existing VLA can adopt.
β Back to CoRL-2025