<a id="top"></a>
<p align="center"><a href="../README.md">&larr; Main README</a> &nbsp;&middot;&nbsp; <a href="awesome-human-centric-ai-survey-resources.md">Our Survey</a></p>

<h1 align="center"><img src="../assets/level-icons/interaction-modeling.png" width="46" height="46" align="absmiddle" alt=""> &nbsp; IV. Interaction Modeling</h1>

<p align="center">Research on how humans relate to objects, scenes, and other people through physical and social behavior.</p>

## Browse Categories

<table>
<tr>
<td width="33%" align="center" valign="middle">
<a href="#human-object-interaction"><strong>IV.1 Human-Object Interaction</strong></a>
</td>
<td width="33%" align="center" valign="middle">
<a href="#human-scene-interaction"><strong>IV.2 Human-Scene Interaction</strong></a>
</td>
<td width="33%" align="center" valign="middle">
<a href="#social-interaction"><strong>IV.3 Social Interaction</strong></a>
</td>
</tr>
</table>

---

<a id="human-object-interaction"></a>

## IV.1 Human-Object Interaction

*Relational modeling of human motion, object state, contact, and affordance.*

| Method | Paper | Venue | Paper Page | Website |
|---|---|:---:|:---:|:---:|
| Single-Query Bimanual HOI | Single-Query Person-Centric Bimanual Hand-Object Interaction Detection | ECCV 2026 | [:page_facing_up:](https://link.springer.com/chapter/10.1007/978-3-032-37016-7_20 "Paper page") | - |
| MILO | Reconstructing Humans and Objects in Interaction using Large Reconstruction Models | ECCV 2026 | [:page_facing_up:](https://arxiv.org/abs/2608.27407 "Paper page") | [:house:](https://ac5113.github.io/MILO/ "Homepage") [:octocat:](https://github.com/ac5113/MILO "GitHub") |
| HOIMask | HOIMask: Towards Generative Masked Modeling for Human Object Interaction Generation | ECCV 2026 | [:page_facing_up:](https://arxiv.org/abs/2608.15141 "Paper page") | [:house:](https://jyhflash.github.io/HOIMask/ "Homepage") |
| EgoTac | EgoTac: In-the-wild Tactile Prediction from Egocentric Vision | arXiv 2026 | [:page_facing_up:](https://arxiv.org/abs/2608.15060 "Paper page") | - |
| EgoPHI | EgoPHI: Estimating Contact and Force from Egocentric Vision | ECCV 2026 | [:page_facing_up:](https://arxiv.org/abs/2608.13014 "Paper page") | [:house:](https://egophii.github.io/ "Homepage") |
| HOPE | HOPE: Hand-Object Pressure Estimation from Monocular Videos | arXiv 2026 | [:page_facing_up:](https://arxiv.org/abs/2608.06192 "Paper page") | [:house:](https://subin6.github.io/page-hope/ "Homepage") |
| StreamHOI | StreamHOI: Interaction-aware Temporal Memory Adaptation for Streaming HOI Video Generation | arXiv 2026 | [:page_facing_up:](https://arxiv.org/abs/2607.20174 "Paper page") | [:octocat:](https://github.com/KlingAIResearch/StreamHOI "GitHub") |
| HOMIE | HOMIE: Human-Object Centric Video Personalization via Multimodal Intelligent Enhancement | arXiv 2026 | [:page_facing_up:](https://arxiv.org/abs/2607.18217 "Paper page") | [:octocat:](https://github.com/YIYANGCAI/HOMIE "GitHub") [:house:](https://yiyangcai.github.io/homie-page.github.io/ "Homepage") |
| Uni-HOI | Uni-HOI: A Unified framework for Learning the Joint distribution of Text and Human-Object Interaction | arXiv 2026 | [:page_facing_up:](https://arxiv.org/pdf/2604.27491 "Paper page") | - |
| CoInteract | CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation | arXiv 2026 | [:page_facing_up:](https://arxiv.org/abs/2604.19636 "Paper page") | [:house:](https://xinxiaozhe12345.github.io/CoInteract_Project "Homepage") |
| OmniShow | OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation | ICML 2026 | [:page_facing_up:](https://arxiv.org/abs/2604.11804 "Paper page") | [:octocat:](https://github.com/Correr-Zhou/OmniShow "GitHub") |
| ViHOI | ViHOI: Human-Object Interaction Synthesis with Visual Priors | CVPR 2026 | [:page_facing_up:](https://arxiv.org/abs/2603.24383 "Paper page") | [:octocat:](https://github.com/MPI-Lab/ViHOI "GitHub") |
| ArtHOI | ArtHOI: Taming Foundation Models for Monocular 4D Reconstruction of Hand-Articulated-Object Interactions | CVPR 2026 | [:page_facing_up:](https://openaccess.thecvf.com/content/CVPR2026/html/Wang_ArtHOI_Taming_Foundation_Models_for_Monocular_4D_Reconstruction_of_Hand-Articulated-Object_CVPR_2026_paper.html "Paper page") | [:house:](https://arthoi-reconstruction.github.io/ "Homepage") |
| HOI-PAGE | HOI-PAGE: Zero-Shot Human-Object Interaction Generation with Part Affordance Guidance | ICML 2026 | [:page_facing_up:](https://arxiv.org/pdf/2506.07209 "Paper page") | [:octocat:](https://github.com/craigleili/HOI-PAGE "GitHub") |
| InterPrior | InterPrior: Scaling generative control for physics-based human-object interactions | CVPR 2026 | [:page_facing_up:](https://openaccess.thecvf.com/content/CVPR2026/html/Xu_InterPrior_Scaling_Generative_Control_for_Physics-Based_Human-Object_Interactions_CVPR_2026_paper.html "Paper page") | [:house:](https://sirui-xu.github.io/InterPrior/ "Homepage") |
| OneHOI | OneHOI: Unifying Human-Object Interaction Generation and Editing | CVPR 2026 | [:page_facing_up:](https://openaccess.thecvf.com/content/CVPR2026/html/Hoe_OneHOI_Unifying_Human-Object_Interaction_Generation_and_Editing_CVPR_2026_paper.html "Paper page") | [:house:](https://jiuntian.github.io/OneHOI/ "Homepage") |
| ReGenHOI | ReGenHOI: Unifying Reconstruction and Generation for 3D Human-Object Interaction Understanding | CVPR 2026 | [:page_facing_up:](https://openaccess.thecvf.com/content/CVPR2026/html/Xu_ReGenHOI_Unifying_Reconstruction_and_Generation_for_3D_Human-Object_Interaction_Understanding_CVPR_2026_paper.html "Paper page") | [:octocat:](https://github.com/xumiao66/ReGenHOI "GitHub") |
| ByteLOOM | ByteLoom: Weaving Geometry-Consistent Human-Object Interactions through Progressive Curriculum Learning | arXiv 2025 | [:page_facing_up:](https://arxiv.org/pdf/2512.22854 "Paper page") | [:house:](https://neutrinoliu.github.io/byteloom/ "Homepage") |
| EasyHOI | EasyHOI: Unleashing the Power of Large Models for Reconstructing Hand-Object Interactions in the Wild | CVPR 2025 | [:page_facing_up:](http://openaccess.thecvf.com/content/CVPR2025/html/Liu_EasyHOI_Unleashing_the_Power_of_Large_Models_for_Reconstructing_Hand-Object_CVPR_2025_paper.html "Paper page") | [:octocat:](https://github.com/lym29/EasyHOI "GitHub") |
| HOIGPT | HOIGPT: Learning Long Sequence Hand-Object Interaction with Language Models | CVPR 2025 | [:page_facing_up:](https://openaccess.thecvf.com/content/CVPR2025/html/Huang_HOIGPT_Learning_Long-Sequence_Hand-Object_Interaction_with_Language_Models_CVPR_2025_paper.html "Paper page") | [:house:](https://www.mingzhenhuang.com/projects/hoigpt.html "Homepage") |
| OpenHOI | OpenHOI: Open-World Hand-Object Interaction Synthesis with Multimodal Large Language Model | NeurIPS 2025 | [:page_facing_up:](https://proceedings.neurips.cc/paper_files/paper/2025/hash/f376f5dff6f6ec6364aea7a46ab49574-Abstract-Conference.html "Paper page") | [:octocat:](https://github.com/Zhenhao-Zhang/OpenHOI "GitHub") |
| PrimHOI | PrimHOI: Compositional Human-Object Interaction via Reusable Primitives | ICCV 2025 | [:page_facing_up:](https://openaccess.thecvf.com/content/ICCV2025/html/Jia_PrimHOI_Compositional_Human-Object_Interaction_via_Reusable_Primitives_ICCV_2025_paper.html "Paper page") | [:octocat:](https://github.com/Kairobo/PrimitiveHOI "GitHub") |
| TriDi | TriDi: Trilateral Diffusion of 3D Humans, Objects, and Interactions | ICCV 2025 | [:page_facing_up:](https://openaccess.thecvf.com/content/ICCV2025/html/Petrov_TriDi_Trilateral_Diffusion_of_3D_Humans_Objects_and_Interactions_ICCV_2025_paper.html "Paper page") | [:octocat:](https://github.com/ptrvilya/tridi "GitHub") |

<p align="right"><a href="#top">Back to top &uarr;</a></p>

---

<a id="human-scene-interaction"></a>

## IV.2 Human-Scene Interaction

*Human behavior modeled with the spatial and functional constraints of surrounding scenes.*

| Method | Paper | Venue | Paper Page | Website |
|---|---|:---:|:---:|:---:|
| RESELF | Seeing the World and the Self from Egocentric Video | arXiv 2026 | [:page_facing_up:](https://arxiv.org/abs/2609.01276 "Paper page") | [:house:](https://ka1guan.github.io/RESELF/ "Homepage") [:octocat:](https://github.com/Ka1Guan/RESELF "GitHub") |
| ReViV | ReViV: Reconstructing the Viewer and the View in 4D from Monocular Egocentric Video | ECCV 2026 | [:page_facing_up:](https://arxiv.org/abs/2607.17790 "Paper page") | [:house:](https://reviv4d.github.io/ "Homepage") [:octocat:](https://github.com/lvsean/reviv4d "GitHub") |
| SHOW | Scene and Human in One World: Reconstruction in a Feedforward Pass | arXiv 2026 | [:page_facing_up:](https://arxiv.org/abs/2606.27720 "Paper page") | [:house:](https://bowieshi.github.io/SHOW-project-page/ "Homepage") |
| IMU-to-4D | Seeing Without Eyes: 4D Human-Scene Understanding from Wearable IMUs | arXiv 2026 | [:page_facing_up:](https://arxiv.org/abs/2604.21926 "Paper page") | [:house:](https://tianhang-cheng.github.io/IMU4D/ "Homepage") |
| UniCon3R | UniCon3R: Contact-aware 3D Human-Scene Reconstruction from Monocular Video | arXiv 2026 | [:page_facing_up:](https://arxiv.org/pdf/2604.19923 "Paper page") | [:octocat:](https://github.com/surtantheta/UniCon3R "GitHub") |
| InHabit | InHabit: Leveraging Image Foundation Models for Scalable 3D Human Placement | arXiv 2026 | [:page_facing_up:](https://arxiv.org/abs/2604.19673 "Paper page") | [:octocat:](https://github.com/nibox/Inhabit "GitHub") |
| FunHSI | Open-Vocabulary Functional 3D Human-Scene Interaction Generation | arXiv 2026 | [:page_facing_up:](https://arxiv.org/abs/2601.20835 "Paper page") | [:house:](https://jliu4ai.github.io/projects/funhsi/ "Homepage") |
| UniSH | UniSH: Unifying Scene and Human Reconstruction in a Feed-Forward Pass | arXiv 2026 | [:page_facing_up:](https://arxiv.org/abs/2601.01222 "Paper page") | [:octocat:](https://github.com/murphylmf/UniSH "GitHub") |
| HSI-GPT2 | HSI-GPT2: A Dual-Granularity Large Motion Reasoning Model with Diffusion Refinement for Human-Scene Interaction | CVPR 2026 | [:page_facing_up:](https://openaccess.thecvf.com/content/CVPR2026/papers/Wang_HSI-GPT2_A_Dual-Granularity_Large_Motion_Reasoning_Model_with_Diffusion_Refinement_CVPR_2026_paper.pdf "Paper page") | - |
| Human3R | Human3r: Everyone everywhere all at once | ICLR 2026 | [:page_facing_up:](https://arxiv.org/pdf/2510.06219 "Paper page") | [:house:](https://fanegg.github.io/Human3R/ "Homepage") |
| Uni-Inter | Uni-Inter: Unifying 3D Human Motion Synthesis Across Diverse Interaction Contexts | SIGGRAPH Asia 2025 | [:page_facing_up:](https://arxiv.org/pdf/2511.13032 "Paper page") | - |
| HIS-GPT | His-gpt: Towards 3d human-in-scene multimodal understanding | ICCV 2025 | [:page_facing_up:](https://arxiv.org/pdf/2503.12955 "Paper page") | [:octocat:](https://github.com/ZJHTerry18/HumanInScene "GitHub") |
| HSI-GPT | HSI-GPT: A General-Purpose Large Scene-Motion-Language Model for Human Scene Interaction | CVPR 2025 | [:page_facing_up:](https://openaccess.thecvf.com/content/CVPR2025/html/Wang_HSI-GPT_A_General-Purpose_Large_Scene-Motion-Language_Model_for_Human_Scene_Interaction_CVPR_2025_paper.html "Paper page") | - |
| TRUMANS | Scaling Up Dynamic Human-Scene Interaction Modeling | CVPR 2024 | [:page_facing_up:](https://openaccess.thecvf.com/content/CVPR2024/html/Jiang_Scaling_Up_Dynamic_Human-Scene_Interaction_Modeling_CVPR_2024_paper.html "Paper page") | [:octocat:](https://github.com/jnnan/trumans_utils "GitHub") |

<p align="right"><a href="#top">Back to top &uarr;</a></p>

---

<a id="social-interaction"></a>

## IV.3 Social Interaction

*Reciprocal, participant-aware modeling of interpersonal behavior and communication.*

| Method | Paper | Venue | Paper Page | Website |
|---|---|:---:|:---:|:---:|
| Motion-Omni | Motion-Omni: End-to-End Joint Speech and Full-Body Motion for Spoken Dialogue | arXiv 2026 | [:page_facing_up:](https://arxiv.org/abs/2609.04250 "Paper page") | [:house:](https://step-out.github.io/Motion-Omni-Page/ "Homepage") [:octocat:](https://github.com/step-out/Motion-Omni "GitHub") [🤗](https://huggingface.co/datasets/ChengqianMa/Motion-Omni "Hugging Face") |
| Super Star | Super Star: Towards Streaming Real-time Interactive Agents for Digital Humans | ACM MM 2026 | [:page_facing_up:](https://arxiv.org/abs/2608.24909 "Paper page") | [:house:](https://super-star-2026.github.io/ "Homepage") [:octocat:](https://github.com/PeterIverson/Super-Star "GitHub") |
| OmniMate | OmniMate: Open-Ended Real-Time Streaming Audio-Visual Generation for Interactive Avatars | arXiv 2026 | [:page_facing_up:](https://arxiv.org/abs/2607.23023 "Paper page") | - |
| SocialStructureHHI | Social Structure Matters in 3D Human-Human Interaction Generation | arXiv 2026 | [:page_facing_up:](https://arxiv.org/abs/2606.24255 "Paper page") | [:octocat:](https://github.com/EngineeringAI-LAB/SocialStructureHHI "GitHub") |
| DyaPlex | DyaPlex: Full-Duplex Speech-Motion Model for Dyadic Interaction | arXiv 2026 | [:page_facing_up:](https://arxiv.org/abs/2606.03874 "Paper page") | [:house:](https://research.nvidia.com/labs/amri/projects/DyaPlex/ "Homepage") |
| SocialDirector | SocialDirector: Training-Free Social Interaction Control for Multi-Person Video Generation | arXiv 2026 | [:page_facing_up:](https://arxiv.org/pdf/2605.10079 "Paper page") | - |
| Omni-MMSI-R | Omni-MMSI: Toward Identity-attributed Social Interaction Understanding | CVPR 2026 | [:page_facing_up:](https://arxiv.org/pdf/2604.00267 "Paper page") | [:house:](https://sampson-lee.github.io/omni-mmsi-project-page/ "Homepage") |
| HumanOmni-Speaker | HumanOmni-Speaker: Identifying Who said What and When | arXiv 2026 | [:page_facing_up:](https://arxiv.org/abs/2603.21664 "Paper page") | - |
| ViBES | Vibes: A conversational agent with behaviorally-intelligent 3d virtual body | CVPR 2026 | [:page_facing_up:](https://arxiv.org/pdf/2512.14234 "Paper page") | [:house:](https://ai.stanford.edu/~juze/ViBES/ "Homepage") |
| Mio | Towards Interactive Intelligence for Digital Humans | arXiv 2025 | [:page_facing_up:](https://arxiv.org/abs/2512.13674 "Paper page") | - |
| X-Streamer | X-streamer: Unified human world modeling with audiovisual interaction | arXiv 2025 | [:page_facing_up:](https://arxiv.org/pdf/2509.21574 "Paper page") | [:house:](https://byteaigc.github.io/X-Streamer/ "Homepage") |
| SOLAMI | SOLAMI: Social Vision-Language-Action Modeling for Immersive Interaction with 3D Autonomous Characters | CVPR 2025 | [:page_facing_up:](https://openaccess.thecvf.com/content/CVPR2025/papers/Jiang_SOLAMI_Social_Vision-Language-Action_Modeling_for_Immersive_Interaction_with_3D_Autonomous_CVPR_2025_paper.pdf "Paper page") | [:house:](https://solami-ai.github.io/ "Homepage") |
| VIM | A Unified Framework for Motion Reasoning and Generation in Human Interaction | ICCV 2025 | [:page_facing_up:](https://openaccess.thecvf.com/content/ICCV2025/html/Park_A_Unified_Framework_for_Motion_Reasoning_and_Generation_in_Human_ICCV_2025_paper.html "Paper page") | [:house:](https://vim-motion-language.github.io/ "Homepage") |
| OmniResponse | OmniResponse: Online Multimodal Conversational Response Generation in Dyadic Interactions | NeurIPS 2025 | [:page_facing_up:](https://proceedings.neurips.cc/paper_files/paper/2025/hash/b734c30b9c955c535e333f0301f5e45c-Abstract-Conference.html "Paper page") | [:octocat:](https://github.com/awakening-ai/OmniResponse "GitHub") |

<p align="right"><a href="#top">Back to top &uarr;</a></p>

---

<p align="center"><a href="kinematic-dynamics.md">&larr; Kinematic Dynamics</a> &nbsp;&middot;&nbsp; <a href="awesome-human-centric-ai-survey-resources.md">All Survey Resources</a> &nbsp;&middot;&nbsp; <a href="world-simulation.md">World Simulation &rarr;</a></p>
