RSS 2026 CoCo InEKF - Heungwoo/research GitHub Wiki
CoCo-InEKF: State Estimation with Learned Contact Covariances in Dynamic, Contact-Rich Scenarios
Venue: RSS 2026 (Sydney, Jul 13–17) · Session: Perception and Estimation · paper #178 Authors: Michael Baumgartner, David Müller, Agon Serifi, Ruben Grandia, Espen Knoop, Markus Gross, Moritz Bächer arXiv: 2605.15122 · program page
Summary compiled from the arXiv paper (v1); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Left: the bipedal robot with IMU, actuator measurements, and predefined contact candidate points on the feet. Right: the pipeline — a learned contact module maps proprioception to per-candidate contact velocity covariances, which feed a differentiable Invariant EKF that fuses IMU and leg odometry into the state estimate.
Problem
Proprioceptive state estimation for legged robots traditionally hinges on binary contact detection plus a stationarity assumption, which breaks under partial contact and directional slippage during highly dynamic motion; prior learned contact detectors (e.g., Lin et al.) need labeled contact data and still output binary states. Pure end-to-end estimators, meanwhile, underperform on accuracy and lose filter consistency.
Method
CoCo-InEKF (ETH Zurich / Disney Research) makes the contact-aided Invariant EKF differentiable by permanently maintaining all contact candidates in the state, and replaces binary contacts with continuous contact velocity covariances predicted per candidate by a lightweight MLP contact module. Training is end-to-end via backpropagation through time (BPTT, horizon H=20, buffer L=128) with a simple body-frame velocity L2 state-error loss — no heuristic contact labels; physical interpretability is not enforced, letting the filter exploit non-physical constraints. An automated farthest-point-sampling procedure selects contact candidates. Experiments run on Lima, a custom 0.84 m, 16.2 kg, 20-DoF bipedal robot with a 600 Hz control loop on an onboard Intel i7.
Results
On simulated dancing (RL motion-tracking policy over 81 retargeted Reallusion sequences), CoCo-InEKF achieves the lowest linear-velocity ATE (RMSE 0.046 m/s vs. 0.121–0.123 for the hybrid learned-contact baselines and 2.675 for heuristic-contact InEKF, which diverges). On contact-rich ground motions (N=10 candidates), it reaches RMSE 0.099 m/s, matching much larger end-to-end SET models (0.096–0.107) at a fraction of the cost — 335K vs. 4.8M parameters and 0.18 ms vs. 3.06 ms network inference — and scaling to 18 candidates improves RMSE to 0.069. Automated candidate selection matches or beats hand-picked placements, NEES analysis shows improved filter consistency, and the estimator supports dancing and full-body ground-interaction motions on the physical robot.
Significance
A hybrid estimator design point: keep the InEKF's structure, invariance, and consistency, but learn a richer-than-binary contact representation end-to-end through the filter itself. Useful context for whole-body humanoid control stacks discussed in Review-Humanoid-VLA.
← Back to RSS 2026 survey · RSS-2026-Papers · Home