ICRA 2026 FD VLA - Heungwoo/research GitHub Wiki

FD-VLA โ€” Force-Distilled VLA for Contact-Rich Manipulation

Venue: ICRA 2026 ยท Authors: Ruiteng Zhao, Wenshuo Wang, Yicheng Ma, Xiaocong Li, Francis E.H. Tay, Marcelo H. Ang Jr., Haiyue Zhu ยท arXiv: 2602.02142 Category: Sensory-augmented VLA (force/tactile) Trend tag: Force awareness without a force sensor

Approach diagram

flowchart LR
  VIS["visual observations"] --> FDM
  STATE["robot state"] --> FDM
  QUERY["learnable query token"] --> FDM["Force Distillation Module<br/>(force-vision-state fusion)"]
  FDM --> FT["predicted force token"]
  REALF["real force signal<br/>(latent, training only)"] -. distillation alignment .-> FT
  FT --> VLM["pretrained VLM<br/>(force token injected)"]
  VLM --> ACT["contact-rich action"]
Loading

Problem

Force sensing enables the fine-grained perception needed for dexterous, contact-rich manipulation, and it is a natural modality to add to a Vision-Language-Action (VLA) model. But physical force/torque sensors are expensive, fragile, and absent on most robots, so a force-aware VLA that depends on them at inference cannot deploy broadly. The question FD-VLA poses: can a VLA gain force awareness without ever reading a physical force sensor at deploy time?

Method

FD-VLA introduces a Force Distillation Module (FDM). The FDM takes a learnable query token, conditioned on visual observations and robot state, and maps it to a predicted force token that is aligned (via distillation) with the latent representation of the actual force signal. Real force is therefore used only as a training-time target for the latent alignment; it is not required at inference.

At inference, the distilled force token is injected into the pretrained VLM, enabling force-aware reasoning while preserving the integrity of the model's vision-language semantics (the language/vision pathway is not disrupted). Beyond removing the sensor dependency, the FDM functions as an additional force-vision-state fusion stage placed before the VLM, which the authors argue improves cross-modal alignment and perception-action robustness in contact-rich settings.

Net effect: force-aware manipulation is delivered through a learned token rather than a hardware channel, so the same policy can run on robots that lack force/torque sensors.

Results

The authors report that, in physical experiments, the distilled force token outperforms direct (real) sensor force measurements as well as other baselines โ€” the counter-intuitive headline of the paper, attributed to the FDM's fusion prior cleaning up noisy raw force into a representation better aligned with the VLM. Specific success-rate numbers and the task suite are not confirmable from the public abstract and are omitted here pending the full text.

Significance

FD-VLA sits in the sensory-augmented VLA family (VLA Architectures review) and the force/tactile thread of contact-rich work (Dexterous Manipulation review, tactile cluster). It is a "privileged-information distillation" design โ€” the privileged signal is force, available in training but distilled away for deployment โ€” closely paralleling the tactile-distillation idea in HapticVLA. The provocative claim that a distilled token beats the true sensor signal, if it holds in the full results, reframes force sensing for VLAs as a learnable representation problem rather than a hardware requirement, lowering the cost and fragility barrier to force-aware manipulation.

Links

Related pages

โ† Back to ICRA-2026

โš ๏ธ **GitHub.com Fallback** โš ๏ธ