<a id="top"></a>
<p align="center"><a href="../README.md">&larr; Main README</a> &nbsp;&middot;&nbsp; <a href="awesome-human-centric-ai-survey-resources.md">Our Survey</a></p>

<h1 align="center"><img src="../assets/level-icons/visual-appearance.png" width="46" height="46" align="absmiddle" alt=""> &nbsp; I. Visual Appearance</h1>

<p align="center">Research on how human appearance is perceived, identified, and controllably synthesized in image and video space.</p>

## Browse Categories

<table>
<tr>
<td width="33%" align="center" valign="middle">
<a href="#generalist-human-perception"><strong>I.1 Generalist Human Perception</strong></a>
</td>
<td width="33%" align="center" valign="middle">
<a href="#discriminative-identity-understanding"><strong>I.2 Discriminative Identity Understanding</strong></a>
</td>
<td width="33%" align="center" valign="middle">
<a href="#controllable-human-generation"><strong>I.3 Controllable Human Generation</strong></a>
</td>
</tr>
</table>

---

<a id="generalist-human-perception"></a>

## I.1 Generalist Human Perception

*Reusable human visual priors and shared interfaces across perception objectives.*

| Method | Paper | Venue | Paper Page | Website |
|---|---|:---:|:---:|:---:|
| THFM | THFM: A Unified Video Foundation Model for 4D Human Perception and Beyond | arXiv 2026 | [:page_facing_up:](https://arxiv.org/abs/2603.25892 "Paper page") | - |
| Sapiens2 | Sapiens2 | ICLR 2026 | [:page_facing_up:](https://openreview.net/forum?id=IVAlYCqdvW "Paper page") | [:octocat:](https://github.com/facebookresearch/sapiens2 "GitHub") |
| UniF^2ace | UniF^2ace: A Unified Fine-grained Face Understanding and Generation Model | ICLR 2026 | [:page_facing_up:](https://arxiv.org/abs/2503.08120 "Paper page") | [:octocat:](https://github.com/tulvgengenr/UniF2ace "GitHub") |
| DAViD | DAViD: Data-efficient and Accurate Vision Models from Synthetic Data | ICCV 2025 | [:page_facing_up:](https://openaccess.thecvf.com/content/ICCV2025/html/Saleh_DAViD_Data-efficient_and_Accurate_Vision_Models_from_Synthetic_Data_ICCV_2025_paper.html "Paper page") | [:octocat:](https://github.com/microsoft/DAViD "GitHub") |
| FaceXFormer | FaceXFormer: A Unified Transformer for Facial Analysis | ICCV 2025 | [:page_facing_up:](https://openaccess.thecvf.com/content/ICCV2025/html/Narayan_FaceXFormer_A_Unified_Transformer_for_Facial_Analysis_ICCV_2025_paper.html "Paper page") | [:octocat:](https://github.com/Kartik-3004/facexformer "GitHub") |
| FSFM | FSFM: A Generalizable Face Security Foundation Model via Self-Supervised Facial Representation Learning | CVPR 2025 | [:page_facing_up:](https://openaccess.thecvf.com/content/CVPR2025/html/Wang_FSFM_A_Generalizable_Face_Security_Foundation_Model_via_Self-Supervised_Facial_CVPR_2025_paper.html "Paper page") | [:octocat:](https://github.com/wolo-wolo/FSFM-CVPR25 "GitHub") |
| GroundingFace | GroundingFace: Fine-Grained Face Understanding via Pixel Grounding Multimodal Large Language Model | CVPR 2025 | [:page_facing_up:](http://openaccess.thecvf.com/content/CVPR2025/html/Han_GroundingFace_Fine-grained_Face_Understanding_via_Pixel_Grounding_Multimodal_Large_Language_CVPR_2025_paper.html "Paper page") | - |
| Hulk | Hulk: A Universal Knowledge Translator for Human-Centric Tasks | TPAMI 2025 | [:page_facing_up:](https://arxiv.org/abs/2312.01697 "Paper page") | [:octocat:](https://github.com/OpenGVLab/Hulk "GitHub") |
| L-Man | L-Man: A Large Multi-modal Model Unifying Human-centric Tasks | AAAI 2025 | [:page_facing_up:](https://ojs.aaai.org/index.php/AAAI/article/view/33206 "Paper page") | - |
| Face-MLLM | Face-MLLM: A Large Face Perception Model | arXiv 2024 | [:page_facing_up:](https://arxiv.org/abs/2410.20717 "Paper page") | - |
| Sapiens | Sapiens: Foundation for Human Vision Models | ECCV 2024 | [:page_facing_up:](https://link.springer.com/chapter/10.1007/978-3-031-73235-5_12 "Paper page") | [:octocat:](https://github.com/facebookresearch/sapiens "GitHub") |
| HAP | HAP: Structure-Aware Masked Image Modeling for Human-Centric Perception | NeurIPS 2023 | [:page_facing_up:](https://proceedings.neurips.cc/paper_files/paper/2023/hash/9ed1c94a6c87276f25ebb65231c86c3e-Abstract-Conference.html "Paper page") | [:octocat:](https://github.com/junkunyuan/HAP "GitHub") |
| HumanBench | HumanBench: Towards General Human-Centric Perception with Projector Assisted Pretraining | CVPR 2023 | [:page_facing_up:](http://openaccess.thecvf.com/content/CVPR2023/html/Tang_HumanBench_Towards_General_Human-Centric_Perception_With_Projector_Assisted_Pretraining_CVPR_2023_paper.html "Paper page") | [:octocat:](https://github.com/OpenGVLab/HumanBench "GitHub") |

<p align="right"><a href="#top">Back to top &uarr;</a></p>

---

<a id="discriminative-identity-understanding"></a>

## I.2 Discriminative Identity Understanding

*Person-level recognition, retrieval, and grounding across observations and query settings.*

| Method | Paper | Venue | Paper Page | Website |
|---|---|:---:|:---:|:---:|
| SapiensID 2.0 | SapiensID 2.0: Aligning Human Recognition Foundation Models with Human Perception | arXiv 2026 | [:page_facing_up:](https://arxiv.org/abs/2608.10497 "Paper page") | - |
| ChatReID | ChatReID: Open-ended Interactive Person Retrieval via Hierarchical Progressive Tuning for Vision Language Models | ICCV 2025 | [:page_facing_up:](https://openaccess.thecvf.com/content/ICCV2025/papers/Niu_ChatReID_Open-ended_Interactive_Person_Retrieval_via_Hierarchical_Progressive_Tuning_for_ICCV_2025_paper.pdf "Paper page") | - |
| GIF | GIF: Generative Inspiration for Face Recognition at Scale | CVPR 2025 | [:page_facing_up:](https://openaccess.thecvf.com/content/CVPR2025/html/Ebrahimi_GIF_Generative_Inspiration_for_Face_Recognition_at_Scale_CVPR_2025_paper.html "Paper page") | - |
| IRM++ | Instruct-reid++: Towards universal purpose instruction-guided person re-identification | TPAMI 2025 | [:page_facing_up:](https://arxiv.org/abs/2405.17790 "Paper page") | [:octocat:](https://github.com/hwz-zju/Instruct-ReID "GitHub") |
| LVFace | LVFace: Progressive Cluster Optimization for Large Vision Models in Face Recognition | ICCV 2025 | [:page_facing_up:](https://openaccess.thecvf.com/content/ICCV2025/html/You_LVFace_Progressive_Cluster_Optimization_for_Large_Vision_Models_in_Face_ICCV_2025_paper.html "Paper page") | [:octocat:](https://github.com/bytedance/LVFace "GitHub") |
| RexSeek | Referring to Any Person | ICCV 2025 | [:page_facing_up:](https://openaccess.thecvf.com/content/ICCV2025/html/Jiang_Referring_to_Any_Person_ICCV_2025_paper.html "Paper page") | [:octocat:](https://github.com/IDEA-Research/RexSeek "GitHub") |
| ReID5o | ReID5o: Achieving Omni Multi-modal Person Re-identification in a Single Model | NeurIPS 2025 | [:page_facing_up:](https://proceedings.neurips.cc/paper_files/paper/2025/hash/4d4dbc5c955c5d273eed12d565335068-Abstract-Conference.html "Paper page") | [:octocat:](https://github.com/Zplusdragon/ReID5o_ORBench "GitHub") |
| SapiensID | SapiensID: Foundation for Human Recognition | CVPR 2025 | [:page_facing_up:](https://openaccess.thecvf.com/content/CVPR2025/html/Kim_SapiensID_Foundation_for_Human_Recognition_CVPR_2025_paper.html "Paper page") | [:octocat:](https://github.com/mk-minchul/sapiensid "GitHub") |
| MMPedestron | When Pedestrian Detection Meets Multi-modal Learning: Generalist Model and Benchmark Dataset | ECCV 2024 | [:page_facing_up:](https://arxiv.org/abs/2407.10125 "Paper page") | [:octocat:](https://github.com/BubblyYi/MMPedestron "GitHub") |
| IRM | Instruct-ReID: A Multi-purpose Person Re-identification Task with Instructions | CVPR 2024 | [:page_facing_up:](https://openaccess.thecvf.com/content/CVPR2024/html/He_Instruct-ReID_A_Multi-purpose_Person_Re-identification_Task_with_Instructions_CVPR_2024_paper.html "Paper page") | [:octocat:](https://github.com/hwz-zju/Instruct-ReID "GitHub") |
| PLIP | PLIP: Language-Image Pre-training for Person Representation Learning | NeurIPS 2024 | [:page_facing_up:](https://proceedings.neurips.cc/paper_files/paper/2024/hash/510ad3018bbdc5b6e3b10646e2e35771-Abstract-Conference.html "Paper page") | [:octocat:](https://github.com/zplusdragon/PLIP "GitHub") |

<p align="right"><a href="#top">Back to top &uarr;</a></p>

---

<a id="controllable-human-generation"></a>

## I.3 Controllable Human Generation

*Human synthesis and editing with structured control over intended appearance changes.*

| Method | Paper | Venue | Paper Page | Website |
|---|---|:---:|:---:|:---:|
| H-SPACE | Human-Centric Image Captioning with Subject-Centered Spatial Understanding | ACM MM 2026 | [:page_facing_up:](https://arxiv.org/abs/2609.08300 "Paper page") | [:octocat:](https://github.com/JHang2020/SPACE-Eval "GitHub") |
| WithEveryone | WithEveryone: Unified Planning and Identity Grounding for Group Image Generation | arXiv 2026 | [:page_facing_up:](https://arxiv.org/abs/2608.20336 "Paper page") | [:house:](https://doby-xu.github.io/WithEveryone/ "Homepage") [:octocat:](https://github.com/doby-xu/WithEveryone "GitHub") |
| InstructVVT | InstructVVT: Instruction-Driven Video Virtual Try-On without Auxiliary Spatial Priors | arXiv 2026 | [:page_facing_up:](https://arxiv.org/abs/2608.14070 "Paper page") | [:house:](https://shaodingbao.github.io/InstructVVT/ "Homepage") [:octocat:](https://github.com/ShaoDingBao/InstructVVT "GitHub") |
| Oxygen-TryOn | Oxygen-TryOn: Fashion-Native Foundation Model for Any-item Virtual Try-On | arXiv 2026 | [:page_facing_up:](https://arxiv.org/abs/2607.21694 "Paper page") | [:house:](https://oxygenvision.github.io/Oxygen-TryOn/ "Homepage") |
| Tstars-Tryon 1.0 | Tstars-Tryon 1.0: Robust and Realistic Virtual Try-On for Diverse Fashion Items | arXiv 2026 | [:page_facing_up:](https://arxiv.org/abs/2604.19748 "Paper page") | [🤗](https://huggingface.co/datasets/TaobaoTmall-AlgorithmProducts/Tstars-VTON "Hugging Face") |
| Archon | Archon: A Unified Multimodal Model for Holistic Digital Human Generation | CVPR 2026 | [:page_facing_up:](https://openaccess.thecvf.com/content/CVPR2026/html/Bao_Archon_A_Unified_Multimodal_Model_for_Holistic_Digital_Human_Generation_CVPR_2026_paper.html "Paper page") | [:house:](https://zju3dv.github.io/archon/ "Homepage") |
| Voost | Voost: A Unified and Scalable Diffusion Transformer for Bidirectional Virtual Try-On and Try-Off | SIGGRAPH Asia 2025 | [:page_facing_up:](https://arxiv.org/abs/2508.04825 "Paper page") | [:house:](https://nxnai.github.io/Voost "Homepage") |
| DreamActor-M1 | DreamActor-M1: Holistic, Expressive and Robust Human Image Animation with Hybrid Guidance | ICCV 2025 | [:page_facing_up:](https://openaccess.thecvf.com/content/ICCV2025/html/Luo_DreamActor-M1_Holistic_Expressive_and_Robust_Human_Image_Animation_with_Hybrid_ICCV_2025_paper.html "Paper page") | [:house:](https://grisoon.github.io/DreamActor-M1/ "Homepage") |
| FoundHand | FoundHand: Large-Scale Domain-Specific Learning for Controllable Hand Image Generation | CVPR 2025 | [:page_facing_up:](https://openaccess.thecvf.com/content/CVPR2025/html/Chen_FoundHand_Large-Scale_Domain-Specific_Learning_for_Controllable_Hand_Image_Generation_CVPR_2025_paper.html "Paper page") | [:octocat:](https://github.com/arthurchen0518/FoundHand "GitHub") |
| Visual Persona | Visual Persona: Foundation Model for Full-Body Human Customization | CVPR 2025 | [:page_facing_up:](https://openaccess.thecvf.com/content/CVPR2025/papers/Nam_Visual_Persona_Foundation_Model_for_Full-Body_Human_Customization_CVPR_2025_paper.pdf "Paper page") | [:octocat:](https://github.com/cvlab-kaist/Visual-Persona "GitHub") |
| CosmicMan | CosmicMan: A Text-to-Image Foundation Model for Humans | CVPR 2024 | [:page_facing_up:](https://openaccess.thecvf.com/content/CVPR2024/html/Li_CosmicMan_A_Text-to-Image_Foundation_Model_for_Humans_CVPR_2024_paper.html "Paper page") | [:octocat:](https://github.com/cosmicman-cvpr2024/CosmicMan "GitHub") |
| UPGPT | UPGPT: Universal Diffusion Model for Person Image Generation, Editing and Pose Transfer | ICCVW 2023 | [:page_facing_up:](https://openaccess.thecvf.com/content/ICCV2023W/CV4Metaverse/html/Cheong_UPGPT_Universal_Diffusion_Model_for_Person_Image_Generation_Editing_and_ICCVW_2023_paper.html "Paper page") | [:octocat:](https://github.com/soon-yau/upgpt "GitHub") |

<p align="right"><a href="#top">Back to top &uarr;</a></p>

---

<p align="center"><a href="awesome-human-centric-ai-survey-resources.md">All Survey Resources</a> &nbsp;&middot;&nbsp; <a href="spatial-geometry.md">Spatial Geometry &rarr;</a></p>
