🌎 English | 🇨🇳 中文 | 📄 Paper [RSS 2026]
This repository builds on unitree_rl_gym to train the Unitree Go2 quadruped with reinforcement learning.
For the IsaacLab-based version, see go2_rl_robotlab.
Follow the step-by-step setup guide in setup.md.
Run the following command to launch training:
python legged_gym/scripts/train.py --task=xxx --headless--task: Required. Options includego2,go2_cts,go2_moe_cts,go2_moe_ng_cts,go2_mcp_cts,go2_ac_moe_cts,go2_dual_moe_cts;go2_moe_ctsis the paper's final version.--headless: Render viewer by default; set totrueto disable rendering for higher throughput.--resume: Resume training from a chosen checkpoint in the logs.--experiment_name: Experiment folder to save/load from.--run_name: Run subfolder name to save/load from.--load_run: Name of the run to load (defaults to the most recent run).--checkpoint: Checkpoint index to load (defaults to the latest file).--num_envs: Number of parallel simulated environments.--seed: Random seed.--max_iterations: Maximum training iterations.--sim_device: Physics simulation device. Use--sim_device=cputo force CPU.--rl_device: RL computation device. Use--rl_device=cputo force CPU.--robogauge: Enable RoboGauge evaluation tool; disabled by default. Evaluation results are saved asresults_{it}.yamlinlogs/{exp_name}/{date}/robogauge_resultsand logged to TensorBoard.--robogauge_port: RoboGauge server port; default is 9973.
RoboGauge evaluation requires a separate server to be started. Refer to the RoboGauge documentation.
Default checkpoint path: logs/<experiment_name>/<date_time>_<run_name>/model_<iteration>.pt
The trained model above was evaluated using the RoboGauge framework via Sim2Sim. The models in the table below are the best models after 150k training steps. All released checkpoints are hosted on Hugging Face: wty-yy/go2_rl_gym_data.
| Model | Score | Tracking | Safety | Quality | Level | Download |
|---|---|---|---|---|---|---|
| go2_moe_cts (Ours) | 0.6713 | 0.6669 | 0.7857 | 0.7392 | 7.85 | ckpt |
| go2_ac_moe_cts | 0.6509 | 0.6442 | 0.7644 | 0.7149 | 7.52 | ckpt |
| go2_mcp_cts | 0.6399 | 0.6355 | 0.7542 | 0.7058 | 7.41 | ckpt |
| go2_moe_ng_cts | 0.6519 | 0.6447 | 0.7639 | 0.7186 | 7.56 | ckpt |
| CTS vanilla | 0.5786 | 0.5755 | 0.7066 | 0.6624 | 6.83 | ckpt |
| HIM | 0.5379 | 0.5453 | 0.6476 | 0.6050 | 6.19 | ckpt |
| DreamWaQ | 0.5054 | 0.5105 | 0.6149 | 0.5730 | 5.74 | ckpt |
In the downloaded ckpt files,
*.ptis used for Python deployment, and*.onnxis used for C++ deployment. The models above were all trained with self-collision disabled. In later tests, we found that enabling self-collision can also achieve strong results; see go2_moe_cts_164k_0.6715 - exported with complete model weights - model_164000.pt.
Visualize policies inside Gym with:
python legged_gym/scripts/play.py --task=xxxNotes
- Play launches on randomized terrain with difficulty between 7 and 9.
- It automatically loads the latest checkpoint inside the experiment folder.
- You can specify another model via
experiment_name,load_run, andcheckpoint, for example:python legged_gym/scripts/play.py --task=go2_moe_cts --num_envs 100 --experiment_name go2_cts_hard_terrain --load_run Mar21_22-54-5-46_ --checkpoint 100000
Play exports the Actor network to logs/{experiment_name}/exported/policies:
policy.pt: TorchScript model for Sim2Sim.policy.onnx: ONNX model for Sim2Real.policy.pkl: Raw weights.
Run policies in the Mujoco simulator:
python deploy/deploy_mujoco/deploy_go2.pyConnect an Xbox-compatible gamepad to enable teleoperation; otherwise, the agent keeps a default forward command.
- Swap the policy: The default checkpoint is
deploy/pre_train/go2/go2_cts_150k.pt. Replacepolicy_pathin the YAML config with your ownlogs/{experiment_name}/exported/policies/policy.pt. - Swap terrains: Default terrain is
resources/robots/go2/stairs.xml. Alternatives includeflat.xml,race_track.xml,cross_stairs.xml, andcross_slope.xml. Generate new terrains with windigal - mujoco_terrains
| Flat | Stairs | Race Track |
|---|---|---|
![]() |
![]() |
![]() |
# Onboard Jetson: pick Python by JetPack version
# JetPack 6: Python 3.10
# JetPack 5: Python 3.8
conda create -n deploy python=3.10
conda activate deploy
# Install the matching PyTorch wheel for your Jetson
# https://forums.developer.nvidia.com/t/pytorch-for-jetson/72048
git clone https://github.com/unitreerobotics/unitree_sdk2_python.git
cd unitree_sdk2_python
pip3 install -e .In the Unitree app, open Device → Service, disable mcf/*, and enable the ota_box service.
Assuming the interface to the low-level controller is eth0:
cd deploy/deploy_real
python deploy_real_go2.py eth0Press start to stand and A to engage the controller.
Follow the usage described in unitree_cpp_deploy.
| Python Deploy | C++ Deploy |
|---|---|
![]() |
![]() |
C++ Deployment: Policy 1/2/4 trained by go2_rl_gym, Policy 3 trained by go2_rl_robotlab.
unitree_cpp_deploy_go2_button_prompts.mp4
This repository would not exist without the following open-source projects:
- unitree_rl_gym: Unitree's core RL training framework.
- legged_gym: Base locomotion environment.
- rsl_rl: Reinforcement learning algorithms.
- mujoco: High-performance CPU physics simulator.
- unitree_sdk2_python: Python hardware interface for deployment.
- unitree_sdk2: C++ hardware interface for deployment.
Related publications implemented in this repo:
Contributors:
- @windigal: CTS algorithm reproduction, terrain generation, video editing
- @wertyuilife2: CTS algorithm reproduction
If you find our work helpful, please cite:
@inproceedings{wu2026robogauge,
title={Toward Reliable Sim-to-Real Predictability for MoE-based Robust Quadrupedal Locomotion},
author={Tianyang Wu and Hanwei Guo and Yuhang Wang and Junshu Yang and Xinyang Sui and Jiayi Xie and Xingyu Chen and Zeyang Liu and Xuguang Lan},
booktitle={Proceedings of Robotics: Science and Systems},
year={2026}
}New contributions follow the MIT License; the original unitree_rl_gym remains under the BSD 3-Clause License.
See the complete LICENSE file for details.







