RSS 2026 More with LESS Local Scene - Heungwoo/research GitHub Wiki
More with LESS – Local Scene Representations for Tactile Imaging
Venue: RSS 2026 (Sydney, Jul 13–17) · Session: Perception and Estimation · paper #167 Authors: Zohar Rimon, Elisei Shafer, Tal Tepper, Daniel Kozin, Alon Malka, Roy Holland, Aviv Tamar arXiv: 2606.14344 · program page
Summary compiled from the arXiv paper (v1); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Figure 1. (a) An automated robotic setup collects the largest soft-body tactile-interaction dataset to date (~800 hours) paired with MRI ground truth of modular phantoms. (b) The LESS architecture represents a tactile scene by capturing local tactile-interaction sequences and generating local visualizations. (c) This yields zero-shot compositional generalization to unseen phantoms with multiple inclusions and different shapes. (d) A proof-of-concept real-time hand-held tactile-imaging system visualizes internal structure.
Problem
Tactile imaging reconstructs the internal structure of soft objects through touch, with uses in medical diagnosis and robotic manipulation. Prior self-supervised approaches represent a scene with a single global, unstructured vector and require robot-controlled sensing — limiting generalization (they would need training on all object combinations) and precluding hand-held use.
Method
The authors propose LESS (Local Encoder for Spatial Sensing), an object-centric representation that exploits the local nature of touch: the tactile response at a position is nearly independent of far-away structure. The scene is modeled as a grid of recurrent-encoder "particles," each with a local receptive field, whose latent states are fused by an image generator into 2D or 3D reconstructions of internal structure. This compositional design generalizes from single-inclusion training phantoms to objects with multiple inclusions and varying sizes, and the local structure supports spatial uncertainty estimation to guide where to palpate next. LESS also enables hand-held imaging via external pose tracking and human-like palpation data (sensor: Xela uSkin), and extends tactile imaging to full 3D.
Results
On Table II (GLOBAL baseline vs. LESS), for in-distribution single inclusions GLOBAL is marginally better (F1 85.2% vs 83.5%; diameter error 1.8 vs 1.5 mm). LESS generalizes far better out of distribution: on multiple inclusions it reaches F1 71.1% vs 44.6%, with area error 46.1 vs 78.4 mm² and center-of-mass error 4.0 vs 7.4 mm; on the large phantom it reaches F1 71.3% vs 1.1%, with CoM error 2.3 vs 26.8 mm and diameter error 2.7 vs 9.0 mm. A receptive-field size of β ≈ 10 mm gives the best overall trade-off.
Significance
Brings object-centric, compositional representation learning to soft-object tactile perception and demonstrates the first SSL-based hand-held tactile imager, extending themes in Review-Tactile-VLA and Review-Independent-Visual-Representation.
← Back to RSS 2026 survey · RSS-2026-Papers · Home