Review Demystifying Action Space - Heungwoo/research GitHub Wiki

Demystifying Action Space Design โ€” In-Depth Review (EEF vs Joint, Absolute vs Delta)

Venue: ICML 2026 ยท arXiv: 2602.23408 Affiliations: Tsinghua University (IAIR / Wuxi Research Institute of Applied Technologies) Scale: 13,000+ real-world rollouts ยท 500+ trained models ยท 14 tasks (4 real + 10 sim) ยท 3 platforms One-pager: [ICML-2026-Demystifying-Action-Space]] ยท See also [VLA Architectures

TL;DR

This is the large-scale empirical study that asks: for an imitation-learning manipulation policy, should the action be in joint space or end-effector (task) space, and should it be absolute or delta? The answer, from 13k+ real rollouts, is not "one wins" โ€” it is a decision rule:

  • Temporal axis (absolute vs delta): delta wins, almost always โ€” but only if implemented chunk-wise, not step-wise.
  • Spatial axis (joint vs EEF): complementary. Joint space โ†’ control stability and gets better with scale; task/EEF space โ†’ generalization (cross-embodiment, transfer).
  • Practical rule: single platform, maximize performance โ†’ joint + chunk-wise delta. Cross-embodiment / transfer โ†’ task-space (EEF).

Overview of the action-space study (Figure 1 from Yu et al., 2026)


1. Why this study

Modern VLA / IL work pours effort into scaling data and model size, but the action space โ€” the very thing the policy regresses to โ€” is still chosen by "ad-hoc heuristics or legacy designs" (RT-1 used discretized EEF; ฯ€ uses joint deltas; ACT uses joint absolute; etc.). The paper argues the action space "fundamentally shapes the optimization landscape," and sets out to measure it under controlled conditions instead of folklore.

2. The action-abstraction taxonomy (ยง2)

The design space is factored into two orthogonal axes plus an implementation detail:

Action-space hierarchy and abstraction taxonomy (Figure 2 from Yu et al., 2026)

Spatial abstraction โ€” where the action lives

  • Actuator space (torque / current) โ€” the low-level controller layer.
  • Joint space (configuration, โ„โฟ) โ€” joint angles; reached from EEF via inverse kinematics (risk: singularities).
  • Task space (end-effector SE(3) pose) โ€” the gripper pose; reached from joints via forward kinematics.

Temporal abstraction โ€” how the action is referenced in time

  • Absolute (0th order) โ€” predict the global target position.
  • Delta (1st order) โ€” predict the relative displacement from a reference (needs integration to execute).
  • Higher-order (accel / force) โ€” requires accurate inertial modeling and high-frequency feedback (out of scope for standard IL).

Action chunking โ€” the decisive implementation nuance

When a policy predicts a chunk of future actions, a delta can be referenced two ways:

  • Step-wise delta โ€” each step relative to the previous predicted step (errors integrate along the chunk).
  • Chunk-wise delta โ€” every step relative to the single anchor frame at the chunk's start (errors stay independent).

3. Experimental setup (ยง3)

  • Policies: ACT (regression backbone) and DP (flow-matching / diffusion policy) โ€” so conclusions aren't tied to one model class.
  • Platforms: single-arm AgileX, bimanual AgileX, and RoboTwin 2.0 simulation.
  • Tasks/data: 14 tasks (4 real + 10 sim); 250 demos/task (real), 50/task (sim); trained to 600 epochs (and 900/1200 for scaling).
  • Scale knobs: data {100, 250, 500} trajectories ร— compute {600, 900, 1200} epochs; single- and multi-task; transfer from ฯ€0; cross-embodiment.

4. Findings

RQ1 โ€” Implementation nuances are decisive: chunk-wise โ‰ซ step-wise delta

Chunk-wise delta consistently and significantly beats step-wise delta โ€” the gap reaches upwards of 10% on average. The paper gives the mechanism as a stability theorem:

Proposition 4.1 (noise amplification). For a predicted chunk of length k with bounded per-step noise, the decoded execution error scales as O(k) for step-wise delta (the transform is the lower-triangular all-ones matrix Lโ‚–, so errors accumulate along the horizon), but stays O(1) for chunk-wise delta and absolute actions (the transform is the identity Iโ‚–, so errors propagate independently).

So step-wise delta structurally amplifies prediction noise as the horizon grows; chunk-wise does not. Takeaway: if you use delta actions, always anchor them chunk-wise.

Horizon coupling: absolute control prefers a longer execution horizon; delta control peaks at a shorter horizon (relative reps are sensitive to execution drift). They train at k=60 (2 s @ 30 Hz) and grid-search the execution window 15โ€“60 โ€” i.e., the chunk horizon is not a constant, it must be tuned to the temporal abstraction.

Chunk-wise vs step-wise delta, and the horizon grid-search (Figure 3 from Yu et al., 2026)

RQ2 โ€” Systematic trends

Temporal (delta vs absolute): with the best implementation of each, delta still consistently and significantly outperforms absolute across all platforms, tasks, and both ACT and DP. Why: (1) mapping high-dim vision โ†’ global coordinates has low local coherence, whereas predicting immediate displacement is a more tractable inductive bias; (2) absolute reps need long horizons that are hard to train.

Spatial (joint vs EEF) โ€” the headline comparison: the two are complementary, not ranked:

  • Joint space โ†’ the more robust / stable spatial representation in most single-platform cases.
  • Task space (EEF) โ†’ competitive in low-data/limited-compute regimes and better for generalization.

RQ3 โ€” Consistency under scaling, transfer, cross-embodiment

  • Delta stays superior as data and compute scale (sometimes by only a marginal margin).
  • Joint-space superiority becomes more pronounced as epochs and data increase (especially for regression policies) โ€” it "benefits disproportionately from stronger modeling and extensive training," apparently because it better captures the underlying kinematic manifold.
  • Task-space (EEF) is competitive when data/compute is limited, and โ€” critically โ€” EEF becomes the superior choice under transfer learning (from ฯ€0) and cross-embodiment, where a body-agnostic SE(3) pose transfers across different kinematics while joint vectors do not.

Consistency of action-space superiority under data/compute scaling (Figure 5 from Yu et al., 2026)

5. The EEF-vs-Joint verdict (what you actually came for)

Joint space (โ„โฟ) Task / EEF space (SE(3))
Strength Control stability; captures the kinematic manifold Generalization / transfer
Scaling Improves with more data/compute (esp. regression) Strong in low-data / limited-compute
Transfer & cross-embodiment Weaker (joint vectors are body-specific) Superior (SE(3) pose is body-agnostic)
Risks needs IK; singularities FK is well-defined
Best when one hardware platform, maximize performance many bodies, transfer, generalization

And on the temporal axis, delta > absolute in essentially all standard settings โ€” provided it is chunk-wise.

6. Practical guidelines (the paper's conclusion)

  1. The chunking execution horizon k is not a constant โ€” adapt it to the temporal abstraction (longer for absolute, shorter for delta).
  2. Single-platform, resource-rich, maximize performance โ†’ joint space + chunk-wise delta is the most robust combination.
  3. Generalized setting (cross-embodiment / transfer learning) โ†’ task space (EEF) is the superior spatial abstraction.

7. Significance

Most VLA papers inherit an action space without justifying it; this study turns that choice into a measured, reproducible decision and supplies a mechanistic reason (the O(k) vs O(1) noise-propagation result) for the single most common silent bug โ€” step-wise delta chunking. It pairs naturally with VLA Architectures (the action-decoder axis) and the ฯ€-series choice of chunk-wise joint deltas, and it explains why cross-embodiment frameworks (Cross-Embodiment) gravitate to EEF/SE(3) action spaces. As an ICML-style "demystifying" study it is evidence, not a new model โ€” but it is the reference to cite when defending an action-space choice.

No public code repository was found at review time; analysis is grounded in the arXiv paper (v1).

Links

Related pages

โ† Back to ICML-2026 ยท Home