Review Demystifying Action Space - Heungwoo/research GitHub Wiki
Demystifying Action Space Design โ In-Depth Review (EEF vs Joint, Absolute vs Delta)
Venue: ICML 2026 ยท arXiv: 2602.23408 Affiliations: Tsinghua University (IAIR / Wuxi Research Institute of Applied Technologies) Scale: 13,000+ real-world rollouts ยท 500+ trained models ยท 14 tasks (4 real + 10 sim) ยท 3 platforms One-pager: [ICML-2026-Demystifying-Action-Space]] ยท See also [VLA Architectures
TL;DR
This is the large-scale empirical study that asks: for an imitation-learning manipulation policy, should the action be in joint space or end-effector (task) space, and should it be absolute or delta? The answer, from 13k+ real rollouts, is not "one wins" โ it is a decision rule:
- Temporal axis (absolute vs delta): delta wins, almost always โ but only if implemented chunk-wise, not step-wise.
- Spatial axis (joint vs EEF): complementary. Joint space โ control stability and gets better with scale; task/EEF space โ generalization (cross-embodiment, transfer).
- Practical rule: single platform, maximize performance โ joint + chunk-wise delta. Cross-embodiment / transfer โ task-space (EEF).

1. Why this study
Modern VLA / IL work pours effort into scaling data and model size, but the action space โ the very thing the policy regresses to โ is still chosen by "ad-hoc heuristics or legacy designs" (RT-1 used discretized EEF; ฯ uses joint deltas; ACT uses joint absolute; etc.). The paper argues the action space "fundamentally shapes the optimization landscape," and sets out to measure it under controlled conditions instead of folklore.
2. The action-abstraction taxonomy (ยง2)
The design space is factored into two orthogonal axes plus an implementation detail:

Spatial abstraction โ where the action lives
- Actuator space (torque / current) โ the low-level controller layer.
- Joint space (configuration, โโฟ) โ joint angles; reached from EEF via inverse kinematics (risk: singularities).
- Task space (end-effector SE(3) pose) โ the gripper pose; reached from joints via forward kinematics.
Temporal abstraction โ how the action is referenced in time
- Absolute (0th order) โ predict the global target position.
- Delta (1st order) โ predict the relative displacement from a reference (needs integration to execute).
- Higher-order (accel / force) โ requires accurate inertial modeling and high-frequency feedback (out of scope for standard IL).
Action chunking โ the decisive implementation nuance
When a policy predicts a chunk of future actions, a delta can be referenced two ways:
- Step-wise delta โ each step relative to the previous predicted step (errors integrate along the chunk).
- Chunk-wise delta โ every step relative to the single anchor frame at the chunk's start (errors stay independent).
3. Experimental setup (ยง3)
- Policies: ACT (regression backbone) and DP (flow-matching / diffusion policy) โ so conclusions aren't tied to one model class.
- Platforms: single-arm AgileX, bimanual AgileX, and RoboTwin 2.0 simulation.
- Tasks/data: 14 tasks (4 real + 10 sim); 250 demos/task (real), 50/task (sim); trained to 600 epochs (and 900/1200 for scaling).
- Scale knobs: data {100, 250, 500} trajectories ร compute {600, 900, 1200} epochs; single- and multi-task; transfer from ฯ0; cross-embodiment.
4. Findings
RQ1 โ Implementation nuances are decisive: chunk-wise โซ step-wise delta
Chunk-wise delta consistently and significantly beats step-wise delta โ the gap reaches upwards of 10% on average. The paper gives the mechanism as a stability theorem:
Proposition 4.1 (noise amplification). For a predicted chunk of length k with bounded per-step noise, the decoded execution error scales as O(k) for step-wise delta (the transform is the lower-triangular all-ones matrix Lโ, so errors accumulate along the horizon), but stays O(1) for chunk-wise delta and absolute actions (the transform is the identity Iโ, so errors propagate independently).
So step-wise delta structurally amplifies prediction noise as the horizon grows; chunk-wise does not. Takeaway: if you use delta actions, always anchor them chunk-wise.
Horizon coupling: absolute control prefers a longer execution horizon; delta control peaks at a shorter horizon (relative reps are sensitive to execution drift). They train at k=60 (2 s @ 30 Hz) and grid-search the execution window 15โ60 โ i.e., the chunk horizon is not a constant, it must be tuned to the temporal abstraction.

RQ2 โ Systematic trends
Temporal (delta vs absolute): with the best implementation of each, delta still consistently and significantly outperforms absolute across all platforms, tasks, and both ACT and DP. Why: (1) mapping high-dim vision โ global coordinates has low local coherence, whereas predicting immediate displacement is a more tractable inductive bias; (2) absolute reps need long horizons that are hard to train.
Spatial (joint vs EEF) โ the headline comparison: the two are complementary, not ranked:
- Joint space โ the more robust / stable spatial representation in most single-platform cases.
- Task space (EEF) โ competitive in low-data/limited-compute regimes and better for generalization.
RQ3 โ Consistency under scaling, transfer, cross-embodiment
- Delta stays superior as data and compute scale (sometimes by only a marginal margin).
- Joint-space superiority becomes more pronounced as epochs and data increase (especially for regression policies) โ it "benefits disproportionately from stronger modeling and extensive training," apparently because it better captures the underlying kinematic manifold.
- Task-space (EEF) is competitive when data/compute is limited, and โ critically โ EEF becomes the superior choice under transfer learning (from ฯ0) and cross-embodiment, where a body-agnostic SE(3) pose transfers across different kinematics while joint vectors do not.

5. The EEF-vs-Joint verdict (what you actually came for)
| Joint space (โโฟ) | Task / EEF space (SE(3)) | |
|---|---|---|
| Strength | Control stability; captures the kinematic manifold | Generalization / transfer |
| Scaling | Improves with more data/compute (esp. regression) | Strong in low-data / limited-compute |
| Transfer & cross-embodiment | Weaker (joint vectors are body-specific) | Superior (SE(3) pose is body-agnostic) |
| Risks | needs IK; singularities | FK is well-defined |
| Best when | one hardware platform, maximize performance | many bodies, transfer, generalization |
And on the temporal axis, delta > absolute in essentially all standard settings โ provided it is chunk-wise.
6. Practical guidelines (the paper's conclusion)
- The chunking execution horizon k is not a constant โ adapt it to the temporal abstraction (longer for absolute, shorter for delta).
- Single-platform, resource-rich, maximize performance โ
joint space + chunk-wise deltais the most robust combination. - Generalized setting (cross-embodiment / transfer learning) โ
task space (EEF)is the superior spatial abstraction.
7. Significance
Most VLA papers inherit an action space without justifying it; this study turns that choice into a measured, reproducible decision and supplies a mechanistic reason (the O(k) vs O(1) noise-propagation result) for the single most common silent bug โ step-wise delta chunking. It pairs naturally with VLA Architectures (the action-decoder axis) and the ฯ-series choice of chunk-wise joint deltas, and it explains why cross-embodiment frameworks (Cross-Embodiment) gravitate to EEF/SE(3) action spaces. As an ICML-style "demystifying" study it is evidence, not a new model โ but it is the reference to cite when defending an action-space choice.
No public code repository was found at review time; analysis is grounded in the arXiv paper (v1).
Links
- arXiv: 2602.23408
- ICML 2026: https://icml.cc/virtual/2026/poster/62805
- One-page summary: ICML-2026-Demystifying-Action-Space
- Figures 1โ5 reproduced from the paper (Tsinghua IAIR, 2026) for scholarly review; ยฉ the authors.