Adaptive Gait Timing for Fault-Tolerant Quadruped Locomotion

Adaptive Gait Timing for Fault-Tolerant Quadruped Locomotion

Giovanbattista Gravina, Luca Rossini, Carlo Rizzardo, Arturo Laurenzi, Nikos Tsagarakis

3 min læsning10. aug. 2026

Reinforcement Learning for Legged Locomotion

While asymmetric actor-critic training is effective, a large gap between privileged and non-privileged observations can limit the actor’s ability to learn complex behaviors from the reward signal alone. Recent works address this by introducing auxiliary representation alignment mechanisms within single-stage reinforcement learning frameworks.

Method

Our approach leverages asymmetric training to bridge the gap between simulation and reality: the critic has access to explicit fault information, while the actor is trained to approximate a privileged latent representation derived from fault-aware observations using only its proprioceptive history. The resulting policy operates at 50 Hz, providing position targets for the robot’s actuated joints and a gait frequency term.

Robot locomotion experiment illustrating adaptive gait timing

Proposed Approach

We formulate fault-tolerant locomotion as a partially observable control problem. While full system state is available during training, the deployed policy only has access to a subset of observations. This asymmetry motivates the use of an asymmetric actor-critic framework with privileged information.

To bridge the information gap between actor and critic, we introduce an auxiliary latent-alignment objective that encourages the actor’s latent representation to approximate the privileged latent representation during training.

Quadruped gait behavior under degraded actuation

Conclusions

This work presented a deep reinforcement learning approach for fault-tolerant locomotion under actuator power loss. The method combines an asymmetric actor-critic architecture with a latent representation alignment objective, enabling the actor to infer privileged information from a history of proprioceptive observations.

A central design choice is the inclusion of a learnable gait-frequency action, allowing the policy to adapt step timing. We argue that for heavier quadrupeds, fast reactive strategies effective on lighter platforms do not scale well due to tighter actuation limits and stronger dynamic coupling. Adaptive gait timing therefore becomes critical under degraded actuation. By leveraging terrain information together with gait frequency modulation, the policy achieves robust locomotion without predefined faulty-leg strategies or restrictive reward shaping.

Simulation results validate the importance of proprioceptive history and latent alignment. Sim-to-sim and zero-shot sim-to-real transfer on a 68 kg quadruped further demonstrate consistent behavior across simulation and hardware. Future work includes integrating onboard perception for terrain reconstruction, learning unified fault-tolerant locomotion and fall-recovery policies, and exploring fault tolerance in hybrid wheeled-legged systems.

Frequently Asked Questions

What problem does the method address? It addresses fault-tolerant locomotion for quadruped robots experiencing actuator power loss.

How does the actor access information during deployment? The deployed actor uses only a history of proprioceptive observations.

What role does adaptive gait timing play? A learnable gait-frequency action allows the policy to adapt step timing under degraded actuation.

What future work is proposed? Future work includes onboard terrain reconstruction, unified locomotion and fall-recovery policies, and fault tolerance in hybrid wheeled-legged systems.

🍪 Cookie-præferencer

Vi bruger cookies til at måle ydeevne. Privatlivspolitik