LocoWM: High-Precision Locomotion through World-Model-Guided Residual Adaptation

1 University of Chinese Academy of Sciences

2 Institute of Automation, Chinese Academy of Sciences

3 Beijing University of Posts and Telecommunications

4 Beijing Jiaotong University

† Equal contribution.

Abstract

What if a robot could correct a disturbance before the deviation becomes observable?

We present LocoWM, a world-model-guided framework that helps robots maintain precise control during locomotion. A world model predicts the effects of a base controller’s proposed action, and a residual adapter uses these predictions to correct anticipated deviations before they become observable. Two-stage training separates learning to move from learning to refine control, allowing the adapter to focus on precision. Across terrain leveling, acceleration compensation, and push recovery, LocoWM improves control precision and disturbance robustness over end-to-end and reactive residual baselines. Real-world demonstrations show a wheeled quadruped transporting unsecured payloads, with a simulation extension to humanoid tray transport.

LocoWM framework and real-world demonstrations for terrain leveling, acceleration compensation, and push recovery
Figure 1. Left: Real-world transport of unsecured payloads across varied terrains. Center: LocoWM runs in a sequential pipeline: the base controller proposes a locomotion action, the world model predicts future task substates conditioned on this action, and the residual adapter uses these predictions to generate a preactive correction. Right: Real-world demonstrations of (a) terrain leveling, (b) acceleration compensation, and (c) push recovery.

Method

LocoWM: Preactive Control

A base controller, a predictive world model, and a residual adapter work together to correct anticipated deviations.

  1. Move

    The base policy proposes a command-following locomotion action.

  2. Predict

    The world model forecasts future task substates from history and the base action.

  3. Adapt

    The residual adapter corrects anticipated deviations before execution.

Two-stage LocoWM training and inference framework
Two-stage training, one-pass prediction. First learn locomotion and task-substate dynamics; then freeze both modules and train the precision adapter.

Real-world experiments

Real-world demonstrations.

Unsecured payloads, uneven terrain, acceleration, and external pushes on a Unitree Go2-W.

Terrain leveling

One-sided bridge · Cup
One-sided bridge · Cup, second view
Slope · Cup
Slope · Cup, second view
One-sided bridge · Bottle
One-sided bridge · Stacked blocks
Bumps · Bottle
Bumps · Stacked blocks
Slope · Bottle
Slope · Stacked blocks
Wavy terrain · Bottle
Wavy terrain · Stacked blocks

Acceleration compensation

Acceleration compensation · Platform pitch

Push recovery

Push forward · Bottle retained
Push backward · Bottle retained
Push right · Platform recovery
Push left · Platform recovery

Simulation

Simulation experiments.

Task-state traces show how posture changes with terrain, acceleration, and disturbances.

Terrain leveling · Platform posture and acceleration across terrain
Acceleration · Velocity tracking and pitch response
Push recovery · Disturbance response on the base frame

Beyond wheeled locomotion

Same idea, Another embodiment.

LocoWM also supports bimanual tray transport with a Unitree G1 humanoid in simulation, keeping an unsecured payload stable during walking.

94.1%Success Rate

Simulation extension · Appendix D, Table 6.
Compared with 91.0% for ReST-RL.

Citation

BibTeX

@misc{zhao2026locowm,
      title={LocoWM: High-Precision Locomotion through World-Model-Guided Residual Adaptation}, 
      author={Zijie Zhao and Shengqian Chen and Xiaoxu Wang and Han Jiang and Yuanheng Zhu and Dongbin Zhao},
      year={2026},
      eprint={2609.39179},
      archivePrefix={arXiv},
      primaryClass={cs.RO},
      url={https://arxiv.org/abs/2609.39179}, 
}