<a id="top"></a>
<p align="center"><a href="../README.md">&larr; Main README</a> &nbsp;&middot;&nbsp; <a href="awesome-human-centric-ai-survey-resources.md">Our Survey</a></p>

<h1 align="center"><img src="../assets/level-icons/embodied-agency.png" width="46" height="46" align="absmiddle" alt=""> &nbsp; VI. Embodied Agency</h1>

<p align="center">Research on physically executable human-like control and the transfer of human experience to embodied agents.</p>

## Browse Categories

<table>
<tr>
<td width="50%" align="center" valign="middle">
<a href="#generalist-humanoid-control"><strong>VI.1 Generalist Humanoid Control</strong></a>
</td>
<td width="50%" align="center" valign="middle">
<a href="#human-to-agent-skill-transfer"><strong>VI.2 Human-to-Agent Skill Transfer</strong></a>
</td>
</tr>
</table>

---

<a id="generalist-humanoid-control"></a>

## VI.1 Generalist Humanoid Control

*Reusable motor priors and policies for physically executable whole-body behavior.*

| Method | Paper | Venue | Paper Page | Website |
|---|---|:---:|:---:|:---:|
| WholeBodyWAM (Motion Priors) | WholeBodyWAM: Learning Whole-Body World Action Models with Scalable Motion Priors | arXiv 2026 | [:page_facing_up:](https://arxiv.org/abs/2609.18197 "Paper page") | [:house:](https://zbzyjya.github.io/WholeBodyWAM/ "Homepage") |
| WholeBodyWAM (WBC Coordination) | WholeBodyWAM: Generalizing Pre-trained World-Action Priors to Humanoid Loco-Manipulation via WBC-Grounded Coordination | arXiv 2026 | [:page_facing_up:](https://arxiv.org/abs/2609.16644 "Paper page") | [:house:](https://wholebodywam.github.io/ "Homepage") |
| TANGO | TANGO: Humanoid Navigation in Cluttered Environments with a Whole-Body Vision-Language-Action Model | CoRL 2026 | [:page_facing_up:](https://arxiv.org/abs/2609.09158 "Paper page") | [:house:](https://tango-vla.github.io/ "Homepage") |
| HumanCLAW | HumanCLAW: Can Vision-Language Models Act Through a Body? | arXiv 2026 | [:page_facing_up:](https://arxiv.org/abs/2607.27180 "Paper page") | [:octocat:](https://github.com/Human-CLAW/HumanCLAW "GitHub") [:house:](https://human-claw.github.io/ "Homepage") |
| GPC | GPC: Large-Scale Generative Pretraining for Transferable Motor Control | SIGGRAPH 2026 | [:page_facing_up:](https://arxiv.org/abs/2606.29148 "Paper page") | [:house:](https://yi-shi94.github.io/gpc-page/ "Homepage") |
| Humanoid-GPT | Humanoid-GPT: Scaling Data and Structure for Zero-Shot Motion Tracking | CVPR 2026 | [:page_facing_up:](https://arxiv.org/abs/2606.03985 "Paper page") | [:octocat:](https://github.com/GalaxyGeneralRobotics/Humanoid-GPT/ "GitHub") |
| SCRIPT | SCRIPT: Scalable Diffusion Policy with Multi-stage Training for Language-driven Physics-Based Humanoid Control | arXiv 2026 | [:page_facing_up:](https://arxiv.org/abs/2605.22894 "Paper page") | [:house:](https://zhanglele12138.github.io/SCRIPT/ "Homepage") |
| VGHuman | Visually-grounded Humanoid Agents | arXiv 2026 | [:page_facing_up:](https://arxiv.org/abs/2604.08509 "Paper page") | [:octocat:](https://github.com/alvinyh/VGHuman "GitHub") |
| BFM-Zero | BFM-Zero: A Promptable Behavioral Foundation Model for Humanoid Control Using Unsupervised Reinforcement Learning | ICLR 2026 | [:page_facing_up:](https://arxiv.org/abs/2511.04131 "Paper page") | [:octocat:](https://github.com/LeCAR-Lab/BFM-Zero "GitHub") |
| CHOREO | CHOREO: Every Humanoid Skill as a Trajectory | Preprint 2026 | [:page_facing_up:](https://ziyisun85-ops.github.io/Project-choreo/paper.pdf "Paper page") | [:house:](https://ziyisun85-ops.github.io/Project-choreo/ "Homepage") [:octocat:](https://github.com/ziyisun85-ops/Project-choreo "GitHub") |
| RoboPerform | Do You Have Freestyle? Expressive Humanoid Locomotion via Audio Control | CVPR 2026 | [:page_facing_up:](https://openaccess.thecvf.com/content/CVPR2026/html/Li_Do_You_Have_Freestyle_Expressive_Humanoid_Locomotion_via_Audio_Control_CVPR_2026_paper.html "Paper page") | [:octocat:](https://github.com/gentlefress/RoboPerform "GitHub") |
| RoboGhost | From language to locomotion: Retargeting-free humanoid control via motion latent guidance | ICLR 2026 | [:page_facing_up:](https://arxiv.org/abs/2510.14952 "Paper page") | [:octocat:](https://github.com/gentlefress/RoboGhost "GitHub") |
| VLM-RMD | Human-Object Interaction via Automatically Designed VLM-Guided Motion Policy | ICLR 2026 | [:page_facing_up:](https://arxiv.org/abs/2503.18349 "Paper page") | [:house:](https://vlm-rmd.github.io/ "Homepage") |
| CLAIMS | Iterative Closed-Loop Motion Synthesis for Scaling the Capabilities of Humanoid Control | CVPR 2026 | [:page_facing_up:](https://openaccess.thecvf.com/content/CVPR2026/html/Xu_Iterative_Closed-Loop_Motion_Synthesis_for_Scaling_the_Capabilities_of_Humanoid_CVPR_2026_paper.html "Paper page") | [:house:](https://wesleyxu224.github.io/CLAIMS/ "Homepage") |
| SONIC | SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control | Science Robotics 2026 | [:page_facing_up:](https://arxiv.org/abs/2511.07820 "Paper page") | [:octocat:](https://github.com/NVlabs/GR00T-WholeBodyControl "GitHub") |
| WholeBodyVLA | WholeBodyVLA: Towards Unified Latent VLA for Whole-body Loco-manipulation Control | ICLR 2026 | [:page_facing_up:](https://arxiv.org/abs/2512.11047 "Paper page") | [:octocat:](https://github.com/OpenDriveLab/WholebodyVLA "GitHub") [:house:](https://wholebodyvla.github.io/ "Homepage") |
| BiBo | Endowing GPT-4 with a Humanoid Body: Building the Bridge Between Off-the-Shelf VLMs and the Physical World | arXiv 2025 | [:page_facing_up:](https://arxiv.org/abs/2511.00041 "Paper page") | - |
| GMT | GMT: General Motion Tracking for Humanoid Whole-Body Control | arXiv 2025 | [:page_facing_up:](https://arxiv.org/abs/2506.14770 "Paper page") | [:octocat:](https://github.com/zixuan417/humanoid-general-motion-tracking "GitHub") |
| GR00T N1 | GR00T N1: An Open Foundation Model for Generalist Humanoid Robots | arXiv 2025 | [:page_facing_up:](https://arxiv.org/abs/2503.14734 "Paper page") | [:octocat:](https://github.com/NVIDIA/Isaac-GR00T "GitHub") |
| BumbleBee | From Experts to a Generalist: Toward General Whole-Body Control for Humanoid Robots | NeurIPS 2025 | [:page_facing_up:](https://proceedings.neurips.cc/paper_files/paper/2025/hash/d8f9cd02831dac9717c1c00329a49cd0-Abstract-Conference.html "Paper page") | [:house:](https://beingbeyond.github.io/BumbleBee/ "Homepage") |
| HOI-HLI | Human-Object Interaction from Human-Level Instructions | ICCV 2025 | [:page_facing_up:](https://arxiv.org/abs/2406.17840 "Paper page") | [:octocat:](https://github.com/zhenkirito123/hoifhli_release "GitHub") |
| Intermimic | Intermimic: Towards Universal Whole-body Control for Physics-based Human-object Interactions | CVPR 2025 | [:page_facing_up:](https://openaccess.thecvf.com/content/CVPR2025/papers/Xu_InterMimic_Towards_Universal_Whole-Body_Control_for_Physics-Based_Human-Object_Interactions_CVPR_2025_paper.pdf "Paper page") | [:house:](https://sirui-xu.github.io/InterMimic "Homepage") |
| TokenHSI | TokenHSI: Unified Synthesis of Physical Human-Scene Interactions through Task Tokenization | CVPR 2025 | [:page_facing_up:](https://openaccess.thecvf.com/content/CVPR2025/html/Pan_TokenHSI_Unified_Synthesis_of_Physical_Human-Scene_Interactions_through_Task_Tokenization_CVPR_2025_paper.html "Paper page") | [:house:](https://liangpan99.github.io/TokenHSI "Homepage") |
| UniPhys | UniPhys: Unified Planner and Controller with Diffusion for Flexible Physics-Based Character Control | ICCV 2025 | [:page_facing_up:](https://openaccess.thecvf.com/content/ICCV2025/html/Wu_UniPhys_Unified_Planner_and_Controller_with_Diffusion_for_Flexible_Physics-Based_ICCV_2025_paper.html "Paper page") | [:octocat:](https://github.com/wuyan01/UniPhys "GitHub") |
| UniTracker | UniTracker: Learning Universal Whole-Body Motion Tracker for Humanoid Robots | RA-L 2025 | [:page_facing_up:](https://ieeexplore.ieee.org/abstract/document/11513991 "Paper page") | [:house:](https://yinkangning0124.github.io/Humanoid-UniTracker/ "Homepage") |
| MaskedMimic | MaskedMimic: Unified Physics-Based Character Control Through Masked Motion Inpainting | TOG 2024 | [:page_facing_up:](https://dl.acm.org/doi/abs/10.1145/3687951 "Paper page") | [:octocat:](https://github.com/NVlabs/ProtoMotions "GitHub") |
| UniHSI | Unified Human-Scene Interaction via Prompted Chain-of-Contacts | ICLR 2024 | [:page_facing_up:](https://arxiv.org/abs/2309.07918 "Paper page") | [:octocat:](https://github.com/OpenRobotLab/UniHSI "GitHub") |

<p align="right"><a href="#top">Back to top &uarr;</a></p>

---

<a id="human-to-agent-skill-transfer"></a>

## VI.2 Human-to-Agent Skill Transfer

*Human experience transformed into adaptable supervision for embodied agents.*

| Method | Paper | Venue | Paper Page | Website |
|---|---|:---:|:---:|:---:|
| HuRo | HuRo: Robotizing Human Videos for Scalable VLA Pretraining | CoRL 2026 | [:page_facing_up:](https://arxiv.org/abs/2609.10706 "Paper page") | [:house:](https://3587jjh.github.io/HuRo/ "Homepage") [:octocat:](https://github.com/3587jjh/HuRo "GitHub") |
| Zero-WAM | Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization | arXiv 2026 | [:page_facing_up:](https://arxiv.org/abs/2608.26103 "Paper page") | [:house:](https://robbyant-research.github.io/Zero-WAM/ "Homepage") [:octocat:](https://github.com/robbyant-research/Zero-WAM "GitHub") |
| HumanScale | HumanScale: Egocentric Human Video Can Outperform Real-Robot Data for Embodied Pretraining | arXiv 2026 | [:page_facing_up:](https://arxiv.org/abs/2606.20521 "Paper page") | [:octocat:](https://github.com/DAGroup-PKU/HumanNet/ "GitHub") |
| HUG | Human Universal Grasping | arXiv 2026 | [:page_facing_up:](https://arxiv.org/abs/2606.17054 "Paper page") | [:octocat:](https://github.com/kevinywu/HUG "GitHub") |
| LARA | LARA: Latent Action Representation Alignment for Vision-Language-Action Models | ICML 2026 | [:page_facing_up:](https://arxiv.org/abs/2606.07100 "Paper page") | [:octocat:](https://github.com/lmy1001/LARA "GitHub") [:house:](https://lmy1001.github.io/ICML26_LARA/ "Homepage") |
| ActiveMimic | ActiveMimic: Egocentric Video Pretraining with Active Perception | arXiv 2026 | [:page_facing_up:](https://arxiv.org/abs/2606.06194 "Paper page") | [:house:](https://activemimic.github.io/ "Homepage") |
| HumanEgo | HumanEgo: Zero-Shot Robot Learning from Minutes of Human Egocentric Videos | arXiv 2026 | [:page_facing_up:](https://arxiv.org/abs/2605.24934 "Paper page") | [:octocat:](https://github.com/TX-Leo/HumanEgo "GitHub") |
| SUGAR | SUGAR: A Scalable Human-Video-Driven Generalizable Humanoid Loco-Manipulation Learning Framework | arXiv 2026 | [:page_facing_up:](https://arxiv.org/abs/2605.20373 "Paper page") | [:octocat:](https://github.com/tianshuwu/SUGAR "GitHub") |
| Being-H0.7 | Being-H0. 7: A Latent World-Action Model from Egocentric Videos | arXiv 2026 | [:page_facing_up:](https://arxiv.org/abs/2605.00078 "Paper page") | [:octocat:](https://github.com/BeingBeyond/Being-H "GitHub") |
| EgoScale | EgoScale: Scaling Dexterous Manipulation with Diverse Egocentric Human Data | arXiv 2026 | [:page_facing_up:](https://arxiv.org/pdf/2602.16710v1 "Paper page") | [:house:](https://research.nvidia.com/labs/gear/egoscale/ "Homepage") |
| DreamDojo | DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos | ICML 2026 | [:page_facing_up:](https://arxiv.org/abs/2602.06949 "Paper page") | [:octocat:](https://github.com/NVIDIA/DreamDojo "GitHub") |
| ActiveUMI | ActiveUMI: Robotic Manipulation with Active Perception from Robot-Free Human Demonstrations | ICRA 2026 | [:page_facing_up:](https://arxiv.org/abs/2510.01607 "Paper page") | [:house:](https://activeumi.github.io/ "Homepage") |
| Being-H0 | Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos | ICML 2026 | [:page_facing_up:](https://arxiv.org/abs/2507.15597 "Paper page") | [:octocat:](https://github.com/BeingBeyond/Being-H0 "GitHub") |
| VITRA | Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos | ICRA 2026 | [:page_facing_up:](https://arxiv.org/abs/2510.21571 "Paper page") | [:octocat:](https://github.com/microsoft/VITRA/ "GitHub") |
| VIPA-VLA | Spatial-Aware VLA Pretraining through Visual-Physical Alignment from Human Videos | CVPR 2026 | [:page_facing_up:](https://openaccess.thecvf.com/content/CVPR2026/html/Feng_Spatial-Aware_VLA_Pretraining_through_Visual-Physical_Alignment_from_Human_Videos_CVPR_2026_paper.html "Paper page") | [:octocat:](https://github.com/BeingBeyond/VIPA-VLA "GitHub") |
| UniDex | UniDex: A Robot Foundation Suite for Universal Dexterous Hand Control from Egocentric Human Videos | CVPR 2026 | [:page_facing_up:](https://openaccess.thecvf.com/content/CVPR2026/html/Zhang_UniDex_A_Robot_Foundation_Suite_for_Universal_Dexterous_Hand_Control_CVPR_2026_paper.html "Paper page") | [:octocat:](https://github.com/unidex-ai/UniDex "GitHub") |
| Human0 / PHSD | In-N-On: Scaling Egocentric Manipulation with In-the-Wild and On-task Data | arXiv 2025 | [:page_facing_up:](https://arxiv.org/abs/2511.15704 "Paper page") | [:house:](https://xiongyicai.github.io/In-N-On "Homepage") |
| EgoVLA | EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos | arXiv 2025 | [:page_facing_up:](https://arxiv.org/abs/2507.12440 "Paper page") | [:octocat:](https://github.com/RchalYang/EgoVLA_Release "GitHub") |
| EgoMimic | EgoMimic: Scaling Imitation Learning via Egocentric Video | ICRA 2025 | [:page_facing_up:](https://ieeexplore.ieee.org/abstract/document/11127989/ "Paper page") | [:octocat:](https://github.com/SimarKareer/EgoMimic "GitHub") |
| ManipTrans | ManipTrans: Efficient Dexterous Bimanual Manipulation Transfer via Residual Learning | CVPR 2025 | [:page_facing_up:](https://openaccess.thecvf.com/content/CVPR2025/html/Li_ManipTrans_Efficient_Dexterous_Bimanual_Manipulation_Transfer_via_Residual_Learning_CVPR_2025_paper.html "Paper page") | [:octocat:](https://github.com/ManipTrans/ManipTrans.git "GitHub") |
| HR-Align | Mitigating the Human-Robot Domain Discrepancy in Visual Pre-training for Robotic Manipulation | CVPR 2025 | [:page_facing_up:](https://openaccess.thecvf.com/content/CVPR2025/html/Zhou_Mitigating_the_Human-Robot_Domain_Discrepancy_in_Visual_Pre-training_for_Robotic_CVPR_2025_paper.html "Paper page") | [:octocat:](https://github.com/jiaming-zhou/HumanRobotAlign "GitHub") |
| Moto | Moto: Latent Motion Token as the Bridging Language for Learning Robot Manipulation from Videos | ICCV 2025 | [:page_facing_up:](https://openaccess.thecvf.com/content/ICCV2025/html/Chen_Moto_Latent_Motion_Token_as_the_Bridging_Language_for_Learning_ICCV_2025_paper.html "Paper page") | [:octocat:](https://github.com/TencentARC/Moto "GitHub") |
| ZeroMimic | ZeroMimic: Distilling Robotic Manipulation Skills from Web Videos | ICRA 2025 | [:page_facing_up:](https://ieeexplore.ieee.org/abstract/document/11128283/ "Paper page") | [:octocat:](https://github.com/junyaoshi/ZeroMimic "GitHub") |

<p align="right"><a href="#top">Back to top &uarr;</a></p>

---

<p align="center"><a href="world-simulation.md">&larr; World Simulation</a> &nbsp;&middot;&nbsp; <a href="awesome-human-centric-ai-survey-resources.md">All Survey Resources</a></p>
