RSS 2026 Mind Your Steps - Heungwoo/research GitHub Wiki

Mind Your Steps: A General Learning Framework for Accurate Humanoid Foothold Tracking

Venue: RSS 2026 (Sydney, Jul 13–17) · Session: Humanoids · paper #28 Authors: Alessandro Montenegro, Shihao Li, Puze Liu, Alberto Maria Metelli, Jan Peters arXiv: 2606.08253 · program page

Summary compiled from the arXiv paper (v1); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Training and deployment architecture (Figure 1 of arXiv 2606.08253, © the authors)

Figure 1: the policy receives proprioception, previous action, the foothold goal g_t, and gait phase, and outputs joint targets for a PD controller. The modular Goal Generator supplies foothold targets: during training a procedural Goal Sampler generates synthetic targets (green box), while at deployment it is swapped for task-specific planners — simulation planners for stairs/cluttered cone fields (top) or a real-world vision-based marker estimator in a corridor (bottom left).

Problem

Velocity-commanded RL locomotion policies are robust but give no explicit control over foot placement, causing unsafe or imprecise stepping; existing foothold-tracking policies rely on unrealistic observations (binary contact flags, precise base localization), live only in simulation, or are welded into task-specific pipelines. The goal is a general-purpose, standalone 3D foothold-tracking low-level controller ready for real deployment.

Method

A lightweight model-free framework trained with asymmetric PPO (actor 512×256×128) built on LocoMuJoCo. The goal vector encodes next left/right foothold position offsets and yaw (quaternion) in the current stance-foot frame, so targets stay constant through each swing phase — removing the need for base state estimation — and no contact flags are observed. A procedural Goal Sampler generates feasible 3D targets each gait switch (perturbed heading angle α, step length d, yaw offset β, height offset z, plus a "hold-still" mode), with terrain realized by height-adjusting pillars; rewards balance swing/stance foothold tracking, swing-window foot clearance, and knee height. The trained policy pairs with arbitrary high-level planners (vision-based estimators, teleoperation, path planners) and targets the Booster T1 humanoid.

Results

In simulation the Foothold-Tracking policy (πFT) consistently beats a state-of-the-art velocity-tracking baseline (πVT) on goal reaching; on narrow bridges (0.15–0.25 m wide, 2–4 m long) πFT keeps high success where πVT degrades sharply. A frame ablation shows root-frame targets (πFT-R) collapse under injected localization noise while stance-foot-frame targets are unaffected. In a cluttered cone field, πFT with a planner reaches 98% success with 15 cones (0% falls). Stairs (straight and spiral, −10° yaw per step) and ramp experiments quantify success and 3D foot placement error as step height/length grow. On the real Booster T1 with an onboard RealSense D455 detecting ground markers, the policy achieves 93.08% success stepping on designated foothold targets, zero-shot from simulation without external localization.

Significance

Positions accurate foothold control as a modular low-level block that upstream planners — and eventually loco-manipulation stacks — can command directly, complementary to the humanoid whole-body-control thread in Review-Humanoid-VLA.

← Back to RSS 2026 survey · RSS-2026-Papers · Home