Humanoid Badminton:
Learning Dynamic Racket Skills from Limited Human Motion Data
Accepted at CoRL 2026
Method Overview
Overview of the three-stage hierarchical reinforcement learning framework. Task-randomized motion augmentation expands sparse annotated hitting events into executable target-conditioned stroke variations, forming a continuous latent skill space. A high-level planner then composes these skills online according to the observed shuttle state, while a context-conditioned adversarial regularizer encourages natural planner-level skill usage without compromising return performance.
Abstract
High-speed racket sports provide a demanding testbed for humanoid robots, requiring time-critical decisions, precise striking, and dynamic whole-body coordination. In badminton, fast-changing shuttle trajectories require timely contact decisions, while successful returns demand precise racket pose and velocity within a brief contact window and across a broad three-dimensional striking workspace. Human motion data provide valuable priors for such athletic skills, but usable badminton references are limited and imperfect. Direct tracking provides insufficient executable variation for diverse shuttle conditions, while purely task-driven optimization may produce unnatural motion. To address these challenges, we present a three-stage hierarchical reinforcement learning framework for dynamic humanoid badminton. First, task-randomized motion augmentation expands sparse annotated hitting events into executable target-conditioned stroke variations, forming a continuous latent skill space. Second, a high-level planner outputs continuous latent skill codes to compose these skills online according to the observed shuttle state. Third, a context-conditioned adversarial regularizer encourages more natural planner-level skill usage while preserving return performance. When deployed on a real humanoid robot, our system achieves sustained multi-skill rallies with human players, including forehand, backhand, and highly dynamic jump returns. This is the first real-world humanoid racket-sport system to demonstrate multi-skill human-robot rallies including highly dynamic jump returns.
Skill Demonstrations
Real-world demonstrations of the learned forehand, backhand, and highly dynamic jump-return skills across multiple executions.
Forehand
Forehand demo 1
Forehand demo 2
Forehand demo 3
Backhand
Backhand demo 1
Backhand demo 2
Backhand demo 3
Jump Return
Jump-return demo 1
Jump-return demo 2
Jump-return demo 3
Human-Robot Rally Demonstrations
Selected examples of human-robot rallies, illustrating repeated shuttle interception and recovery.
Rally clip 1
Rally clip 2
Rally clip 3
BibTeX
@misc{cui2026humanoidbadmintonlearningdynamic,
title={Humanoid Badminton: Learning Dynamic Racket Skills from Limited Human Motion Data},
author={Jingzhi Cui and Zhexiong Wang and Bangjie Xu and Pengyu Zhao and Youyuan Li and Zhi Su and Peng Ren and Mengdi Xu and Chao Yu and Yi Wu and Luyang Wang and Zhongyu Li},
year={2026},
eprint={2609.31840},
archivePrefix={arXiv},
primaryClass={cs.RO},
url={https://arxiv.org/abs/2609.31840},
}
Acknowledgements
This study was in part supported by the InnoHK initiative of the Innovation and Technology Commission of the Hong Kong Special Administrative Region Government via the Hong Kong Centre for Logistics Robotics. The authors thank Junhao Huang, Yueyang Wang, and Jiamu Qin for their assistance in collecting the badminton motion data.