Xinyang Song

I am currently a four-year Ph.D. student at School of Artificial Intelligence, University of Chinese Academy of Sciences (UCAS) and also the NLPR, Institute of Automation, Chinese Academy of Sciences (CASIA), supervised by Prof. Zhenan Sun. Prior to that, I received my bachelor's degree in Automation from Nanjing University in 2022, advised by Prof. Huaxiong Li.

My research interests revolve around unified generation and understanding and controllable vision generation(image&video). If you are interested in my work, please feel free to reach out for discussions or collaborations!

Email  /  Google Scholar  /  Github

profile photo
Publications
UniAlignment: Semantic Alignment for Unified Image Generation, Understanding, Manipulation and Perception
Xinyang Song, Libin Wang, Weining Wang, Shaozhen Liu, Dandan Zheng, Jingdong Chen, Qi Li, Zhenan Sun
AAAI, 2026
[paper] [code]

UniAlignment, a novel dual-branch diffusion-based framework, enables unified image generation, understanding, manipulation and perception through two forms of semantic alignment: intra-modal semantic alignment and inter-modal semantic alignment.

GroupVideo: Multi-Identity Customized Text-to-Video Generation
Xinyang Song, Libin Wang, Jianxin Sun, Qi Li, Dandan Zheng, Jingdong Chen, Zhenan Sun
IEEE Transactions on Multimedia (TMM), 2025
[paper]

GroupVideo, a novel video dit framework that leverages multiple individual photographs to generate identity-customized video.

Fine-grained Text-to-Image Synthesis with Semantic Refinement
Xinyang Song*, Jianxin Sun*, Yingya Zhang, Libin Wang, Qi Li, Zhenan Sun
ICASSP, 2026 (Oral)
[paper]

SeReDiff, a fine-grained text-to-image generation framework based on semantic refinement.

StyleVideo: One-shot Image-Guided Video Style Transfer
Xinyang Song, Xi Yong, Qi Li, Jianxin Sun, Yunfan Liu, Zhenan Sun
Machine Intelligence Research (MIR), 2025

StyleVideo, a one-shot method that realizes image-guided video style transfer by leveraging diffusion models.

3SGen: Unified Subject, Style, and Structure-Driven Image Generation with Adaptive Task-specific Memory
Xinyang Song, Libin Wang, Weining Wang, Zhiwei Li, Jianxin Sun, Dandan Zheng, Jingdong Chen, Qi Li, Zhenan Sun
arxiv, 2025
[paper]

3SGen, a taskaware unified framework that performs subject-driven, style-driven and structure-driven conditioning modes within a single model.

Review of Talking Face Generation
Xinyang Song, Zhiyuan Yan, Muyi Sun, Linlin Dai, Qi Li, Zhenan Sun
Computer Science, 2023

A Chinese survey on the current status and development trends of talking face generation research.

Honors and Awards
  • Merit Student of UCAS, 2024.
  • National Special Prize of China Education Robot Contest, 2020.
  • Excellent League Member of Nanjing University, 2020.
  • People's Scholarship of Nanjing University, 2019, 2020, 2021.
Internship
  • AntGroup, Bailing, Vision Generation (2024.08-now)
  • Meituan, Advertisement, Algorithm (2023.12-2024.2)

Website Template


© Xinyang Song | Last updated: Mar 30, 2026