RSS 2026 CoRAL - Heungwoo/research GitHub Wiki

CoRAL: Contact-Rich Adaptive LLM-based Control for Robotic Manipulation

Venue: RSS 2026 (Sydney, Jul 13–17) · Session: Manipulation 2 · paper #54 Authors: Berk Cicek, Mert Kaan Er, Ozgur S. Oguz arXiv: 2605.02600 · program page

Summary compiled from the arXiv paper (v2); all numbers quoted from the paper. Trend context: RSS 2026 survey.

CoRAL real-world tasks (Figure 1 of arXiv 2605.02600, © the authors)

Real-world execution of CoRAL on a Franka Emika Panda across the six evaluation tasks: push-and-pick a cutting board, pick-and-place a box, pick-and-place in clutter, push with constant force, flip a box, and flip a box using the wall as an extrinsic contact.

Problem

LLMs/VLMs reason well semantically but lack physical grounding and cannot do adaptive control, so applying them directly to contact- and force-critical manipulation (e.g., pivoting an object against a wall) fails; end-to-end VLAs meanwhile need extensive demonstrations and degrade sharply on such tasks. CoRAL (Bilkent University, LiRA Lab) targets zero-shot contact-rich manipulation without demonstrations.

Method

CoRAL is a modular hierarchy that uses the LLM (GPT-4o) as a cost designer, not a controller: FoundationPose extracts 6-DoF object poses from RGB-D, a VLM supplies semantic physical priors (mass/friction estimates), and the LLM synthesizes an MPPI cost function J0 plus a symbolic contact strategy C0. A neuro-symbolic adaptation loop refines this online — inner loop: high-frequency MPPI re-planning with a reactive feedback augmentation (joint impedance at the hardware tier, MPPI at the trajectory tier); outer loop: the LLM diagnoses failures from interaction feedback, refining both parameters (online system identification) and cost structure. A retrieval-based Memory Unit stores successful (J, C) plans for reuse on recurrent tasks. Evaluation runs in robosuite/MuJoCo (7-DoF Franka Panda, plus two LIBERO tasks) and zero-shot on the real Panda.

Results

In simulation (10 trials/task), CoRAL scores 5/10, 10/10, 10/10, 9/10, 9/10, 7/10 on T1–T6, versus OpenVLA-OFT and π0.5 at 0/10 on T1, T4, T6 (both perfect only on plain pick-and-place) and the L2R planner baseline at 0–5/10 on the contact-critical tasks. Ablations: no pose tracking → 0/10 everywhere; memory lifts T1 from 2/10 to 5/10 and T3 from 9/10 to 10/10; the LLM contact strategy makes planning 83.9% more efficient (32 vs. 199 control steps) on Flip-with-Wall. Online adaptation converges a severely biased world model (2.0 kg assumed vs. 0.25 kg true mass) and cuts VLM prior errors from MAE 0.29 kg/0.14 friction to 0.11 kg/0.06 after four refinement cycles. Zero-shot real-robot results: 4/10 to 10/10 across the six tasks (T2, T3 perfect).

Significance

A strong data point for the modular "foundation model as cost designer + sampling-based control" alternative to end-to-end VLAs on force-critical tasks, with an interpretable failure-recovery loop. Complements the wiki's Review-VLA-Evaluation discussion of where VLAs break, and the extrinsic-contact thread in Review-Dexterous-Manipulation.

← Back to RSS 2026 survey · RSS-2026-Papers · Home