ICML 2026 Learning AttributeAffordance Hierarchies in Hyperbolic - Heungwoo/research GitHub Wiki
Attribute–Affordance Hierarchies in Hyperbolic Space — open-vocabulary 3D affordance grounding via hierarchical concept embeddings
Venue: ICML 2026 (Poster) Category: Reasoning Affiliations: Yuxuan Wang, Tong Li, Yihang Zhu, Guangtao Lyu, Yukuan Min, Chenghao Xu, Jiexi Yan, Xu Yang, Cheng Deng
The paper targets open-vocabulary 3D object affordance grounding (OVAG): localizing affordance regions on 3D objects given either interaction images or textual instructions. The authors argue that existing approaches overlook the connections between local object characteristics (attributes) and their functional capabilities (affordances), which constrains both accuracy and generalization. Their motivating example: "a cup handle affords grasping due to its curved shape and appropriate thickness" — i.e., the affordance is a consequence of local geometric attributes, a relationship prior methods do not model explicitly.
The authors present the Attribute–Affordance Hierarchies (AAH) framework, which explicitly models the hierarchical relationships between object-region attributes and affordances. According to the abstract, the methodology has three core ingredients:
- Hypergraph modeling of local regions — hypergraphs capture the relationships among local object regions, going beyond pairwise links to represent higher-order region-level structure.
- Hyperbolic-space concept embeddings — the attribute and affordance concepts are embedded into hyperbolic space, whose geometry is well suited to representing hierarchical (tree-like) organization between attributes and the affordances they imply.
- Counterfactual attribute samples — the framework incorporates counterfactual attribute samples to strengthen the learned attribute–affordance connections, reinforcing which attributes are causally responsible for an affordance.
By combining visual structural modeling (hypergraphs) with hierarchical conceptual information (hyperbolic embeddings), the method grounds affordance regions from either image or text queries.
flowchart LR
A[3D object + image/text query] --> B[Hypergraph over local regions]
B --> C[Attribute–affordance concepts]
C --> D[Hyperbolic-space embedding<br/>hierarchical organization]
E[Counterfactual attribute samples] --> D
D --> F[Grounded affordance region]
The abstract reports that testing demonstrates improved localization performance through the combined modeling of visual structure and hierarchical conceptual information. No specific numeric metrics are stated in the available source (the abstract from the ICML virtual page); the paper has no arXiv HTML version, so quantitative tables could not be retrieved.
Affordance grounding is a key bridge between perception and manipulation in VLA systems, and open-vocabulary generalization is what lets robots act on objects and instructions unseen at training time. Framing affordances as consequences of hierarchical attribute relationships — and using hyperbolic geometry, which naturally encodes hierarchy — is a distinctive structural inductive bias for the 2026 affordance-grounding landscape.
- ICML 2026: https://icml.cc/virtual/2026/poster/63231
← Back to ICML-2026