ICRA 2026 FD VLA - Heungwoo/research GitHub Wiki
Venue: ICRA 2026 ยท Authors: Ruiteng Zhao, Wenshuo Wang, Yicheng Ma, Xiaocong Li, Francis E.H. Tay, Marcelo H. Ang Jr., Haiyue Zhu ยท arXiv: 2602.02142 Category: Sensory-augmented VLA (force/tactile) Trend tag: Force awareness without a force sensor
flowchart LR
VIS["visual observations"] --> FDM
STATE["robot state"] --> FDM
QUERY["learnable query token"] --> FDM["Force Distillation Module<br/>(force-vision-state fusion)"]
FDM --> FT["predicted force token"]
REALF["real force signal<br/>(latent, training only)"] -. distillation alignment .-> FT
FT --> VLM["pretrained VLM<br/>(force token injected)"]
VLM --> ACT["contact-rich action"]
Force sensing enables the fine-grained perception needed for dexterous, contact-rich manipulation, and it is a natural modality to add to a Vision-Language-Action (VLA) model. But physical force/torque sensors are expensive, fragile, and absent on most robots, so a force-aware VLA that depends on them at inference cannot deploy broadly. The question FD-VLA poses: can a VLA gain force awareness without ever reading a physical force sensor at deploy time?
FD-VLA introduces a Force Distillation Module (FDM). The FDM takes a learnable query token, conditioned on visual observations and robot state, and maps it to a predicted force token that is aligned (via distillation) with the latent representation of the actual force signal. Real force is therefore used only as a training-time target for the latent alignment; it is not required at inference.
At inference, the distilled force token is injected into the pretrained VLM, enabling force-aware reasoning while preserving the integrity of the model's vision-language semantics (the language/vision pathway is not disrupted). Beyond removing the sensor dependency, the FDM functions as an additional force-vision-state fusion stage placed before the VLM, which the authors argue improves cross-modal alignment and perception-action robustness in contact-rich settings.
Net effect: force-aware manipulation is delivered through a learned token rather than a hardware channel, so the same policy can run on robots that lack force/torque sensors.
The authors report that, in physical experiments, the distilled force token outperforms direct (real) sensor force measurements as well as other baselines โ the counter-intuitive headline of the paper, attributed to the FDM's fusion prior cleaning up noisy raw force into a representation better aligned with the VLM. Specific success-rate numbers and the task suite are not confirmable from the public abstract and are omitted here pending the full text.
FD-VLA sits in the sensory-augmented VLA family (VLA Architectures review) and the force/tactile thread of contact-rich work (Dexterous Manipulation review, tactile cluster). It is a "privileged-information distillation" design โ the privileged signal is force, available in training but distilled away for deployment โ closely paralleling the tactile-distillation idea in HapticVLA. The provocative claim that a distilled token beats the true sensor signal, if it holds in the full results, reframes force sensing for VLAs as a learnable representation problem rather than a hardware requirement, lowering the cost and fragility barrier to force-aware manipulation.
- arXiv: 2602.02142 ยท HTML
- Dexterous Manipulation review (tactile/force cluster)
- VLA Architectures review (sensory-augmented category)
- HapticVLA (sibling: distill the contact modality away)
- ICRA 2026 Survey
โ Back to ICRA-2026