ICLR 2026 Demystifying DP - Heungwoo/research GitHub Wiki

Demystifying Diffusion Policies โ€” action memorization & the Action Lookup Table

Venue: ICLR 2026 ยท Authors: Chengyang He, Xu Liu, Gadiel Sznaier Camps, Guillaume Sartoretti, Mac Schwager ยท Paper: arXiv 2505.05787 โ€” Demystifying Diffusion Policies: Action Memorization and Simple Lookup Table Alternatives (May 2025) ยท Category: Empirical study + method ยท Trend tag: What diffusion policies actually learn

Approach diagram

flowchart LR
  Test[Test image] --> Enc[Latent embedding]
  Enc --> NN[Find nearest<br/>TRAINING image in latent space]
  NN --> Recall[Recall associated<br/>training action sequence]
  Recall --> Act[Output action]
  subgraph Hypothesis
    DP[Diffusion Policy<br/>implicitly does this]
  end
  subgraph Proposed
    ALT[Action Lookup Table<br/>contrastive encoder = hash fn<br/>explicit nearest-neighbor lookup<br/>+ OOD flag]
  end
Loading

Problem

Diffusion policies are dexterous and robust from few demonstrations, but why they work is unclear. The paper offers a surprising hypothesis: diffusion policies essentially memorize an action lookup table โ€” and that this memorization is beneficial in the sparse-data regime where there isn't enough density to learn genuine action generalization.

Method

  • Hypothesis. At runtime a diffusion policy finds the closest training image in latent space and recalls the associated training action sequence, giving reactivity without action generalization.
  • Evidence. Systematic empirical tests, including a striking probe: conditioned on wildly out-of-distribution images (cats and dogs), the diffusion policy still emits an action sequence drawn from the training data.
  • Proposed alternative โ€” Action Lookup Table (ALT). A lightweight policy that uses a contrastive image encoder as a hash function to index the closest training action sequence โ€” explicitly performing the computation the diffusion policy implicitly learns. ALT also emits an explicit OOD flag when the runtime image is too far from all training images in latent space, serving as a simple runtime monitor.

Results

For relatively small datasets, ALT matches the performance of a diffusion model while requiring only:

  • 0.0034ร— the inference time, and
  • 0.0085ร— the memory footprint,

enabling much faster closed-loop inference on resource-constrained robots. (Task-level success tables omitted here pending the full paper.)

Significance

A deflationary but actionable reframing: in the small-data regime, the value of diffusion policies may come from memorization + nearest-neighbor recall, not learned generalization. If so, an explicit lookup table is far cheaper and adds a free OOD monitor โ€” directly useful for deployment, and a caution for over-attributing "generalization" to diffusion policies. Pairs naturally with the control-theoretic compounding-error analysis at this venue.

Links

Related pages

โ† Back to ICLR-2026

โš ๏ธ **GitHub.com Fallback** โš ๏ธ