Video Presentation
Our adaptive gait framework enables robust humanoid locomotion across challenging terrains, including stairs up to 30cm and slopes up to 26.5°. The asymmetric actor-critic architecture with contrastive learning provides superior terrain adaptability and sim-to-real transfer capabilities.
Abstract
Reinforcement learning has produced remarkable advances in humanoid locomotion, yet a fundamental dilemma persists for real-world deployment: policies must choose between the robustness of reactive proprioceptive control or the proactivity of complex, fragile perception-driven systems. This paper resolves this dilemma by introducing a paradigm that imbues a purely proprioceptive policy with proactive capabilities, achieving the foresight of perception without its deployment-time costs. Our core contribution is a contrastive learning framework that compels the actor's latent state to encode privileged environmental information from simulation. Crucially, this ``distilled awareness" empowers an adaptive gait clock, allowing the policy to proactively adjust its rhythm based on an inferred understanding of the terrain. This synergy resolves the classic trade-off between rigid, clocked gaits and unstable clock-free policies. We validate our approach with zero-shot sim-to-real transfer to a full-sized humanoid, demonstrating highly robust locomotion over challenging terrains, including 30 cm high steps and 26.5° slopes, proving the effectiveness of our method.
Key Results
Adam Lite successfully climbing 30cm stairs using our adaptive gait framework
Robust locomotion across diverse terrains: 26.5° slopes, multi-slope sequences, uneven terrain, and 15cm steps
Contributions
A novel training framework that uses contrastive learning to distill an awareness of privileged environmental properties into a purely proprioceptive policy, bridging the sim-to-real information gap.
A novel method that resolves the classic trade-off between rigid clocked gaits and inefficient clock-free policies, by using the distilled awareness to intelligently inform an adaptive gait clock.
Comprehensive real-world validation on a full-sized humanoid, demonstrating highly robust, zero-shot sim-to-real locomotion over challenging terrains like 30 cm steps and 26.5° slopes, confirming the practical effectiveness of our approach.
Paper
BibTeX
@misc{lu2025contrastiverepresentationlearningrobust,
title={Contrastive Representation Learning for Robust Sim-to-Real Transfer of Adaptive Humanoid Locomotion},
author={Yidan Lu and Rurui Yang and Qiran Kou and Mengting Chen and Tao Fan and Peter Cui and Yinzhao Dong and Peng Lu},
year={2025},
eprint={2509.12858},
archivePrefix={arXiv},
primaryClass={cs.RO},
url={https://arxiv.org/abs/2509.12858},
}
Acknowledgements
We would like to thank David Yan and Arthur Zhang for their valuable contributions and support throughout this research project.