RSS 2026 Emergent Neural Automaton Policies Learning - Heungwoo/research GitHub Wiki

Emergent Neural Automaton Policies: Learning Symbolic Structure from Visuomotor Trajectories

Venue: RSS 2026 (Sydney, Jul 13–17) · Session: Imitation learning 2 · paper #144 Authors: Yiyuan Pan, Xusheng Luo, Hanjiang Hu, Peiqi Yu, Changliu Liu arXiv: 2603.25903 · program page

Summary compiled from the arXiv paper (v2); all numbers quoted from the paper. Trend context: RSS 2026 survey.

ENAP discovered task structures (Figure 1 of arXiv 2603.25903, © the authors)

Automata discovered without labels: for "Push Tee to the green position" (top left) a cyclic Mealy machine over clustered observation symbols; for "Insert the peg into the box" (top right) an automaton with an autonomous failure-recovery transition (q2→q1 to retry insertion); bottom: a branching automaton for the real-world "place the red can into the red bowl" task, with the observation clusters that define the symbolic alphabet. State semantics (e.g., "Hold") are named by GPT from cluster representatives.

Problem

End-to-end policies lack the structural priors needed for long-horizon reasoning, while classical neuro-symbolic TAMP relies on rigid hand-crafted symbolic priors (e.g., temporal logic predicates). ENAP (CMU Robotics Institute) asks whether a bi-level neuro-symbolic policy can emerge from raw visuomotor demonstrations, with no task-specific labels.

Method

Three stages: (1) adaptive symbol abstraction — HDBSCAN clusters encoded trajectory segments into a discrete alphabet, with an RNN history encoder trained via prediction + contrastive losses; (2) structure extraction via an extended L* algorithm that, lacking an interactive oracle, answers membership queries by mining the demonstration dataset, producing a Probabilistic Mealy Machine (PMM) enforcing table closedness/consistency, refined in an EM loop with the policy; (3) a bi-level controller where the PMM supplies a coarse action prior per transition and a low-level residual network (behavior cloning) adds the precise correction, ât = a_base + Δa.

Results

On complex manipulation (DualStackCube / PegInsert), ENAP with a DINOv2 encoder reaches 76.0%/63.2% at 22.94M parameters versus π0 (73.4%/51.6% at 3.3B), OpenVLA (69.8%/42.3% at 7.7B), and Diffusion Policy (41.2%/31.1%); with the oracle encoder it matches the 2.98M-parameter oracle policy (98.8%/85.6%). It stays robust down to 25% of training data, where baselines degrade sharply (the paper claims up to 27% advantage over VLAs in low-data regimes, and at least 8% consistently). On long-horizon TAMP suites, distilling structure from FLOWER's latent tokens lifts sequential 5/5 completion from 90.6% to 96.8% and hierarchical 5/5 from 15.9% to 28.2%. On three real-robot tasks, ENAP (DINO) beats fine-tuned π0.5 (88.24/94.12/94.12 vs. 58.82/76.47/64.71) at 23M vs. 3.4B parameters and 281 ms vs. 6,841 ms inference. The learned automata exhibit interpretable branching and autonomous failure recovery via structural loops.

Significance

Shows that interpretable System-2-like discrete structure can be mined unsupervisedly from demonstrations — even from a pretrained VLA's own token space — and that a tiny automaton-guided policy can beat billion-parameter VLAs in low-data, long-horizon settings. Connects to the wiki's Review-System-0-1-2 hierarchy thread and the data-efficiency debate in Review-LBM-Cotraining.

← Back to RSS 2026 survey · RSS-2026-Papers · Home