|
Xiaoyang Guo (郭晓阳)
I work on embodied foundation models, with a focus on world modeling, robot learning, and 3D perception.
Previously, I was a Principal Researcher at Horizon Robotics, leading research on
3D reconstruction and interactive driving simulation. I received my Ph.D. from
CUHK MMLab, advised by
Prof. Xiaogang Wang and
Prof. Hongsheng Li, and my B.S. in Computer Science from
Tsinghua University.
I am looking for research interns interested in embodied AI, 3D foundation models,
and generative simulation. Please reach out by email.
Email /
CV /
Scholar /
Github
|
|
News
|
- 2026.09: Released XPACE and IronMind, two technical reports on embodied world models and humanoid manipulation.
- 2026.06: One paper accepted to ICML 2026 (EPS3D). SCOPE released on arXiv.
- 2026.04: Four papers accepted to CVPR 2026: Scal3R (Highlight), LongStream, LiteVGGT, and VGGT4D.
- 2026.03: CompoSIA accepted to ECCV 2026. Started working on embodied foundation models.
- 2026.02: SAIL-Recon accepted as an Oral at 3DV 2026.
- 2025.10: Three papers at ICCV 2025 (Epona, DM-Calib, StableDepth Highlight); RAD at NeurIPS 2025; SynthDrive and ComDrive at IROS 2025.
|
Selected Publications
* equal contribution; † project lead.
Full list on Google Scholar.
|
|
|
IronMind: Scaling Humanoid Dexterous Manipulation via Camera-Space Ego-Centric Pretraining
Huimin Pan,
Yufan Ren,
Kunpeng Song,
Siyang Wang,
Xiwen Zhang,
Xiaoyun Hu,
Zhuoxu Duan,
Hanrui Zheng,
Jialeng Ni,
Nathan Zhao,
Sibo Ma,
Zhenxuan Fan,
Zhongyang Che,
Danny Bao,
Jiacheng Wei,
Jerry Bai,
Xiaoyu Yue,
Xiaoyang Guo,
Chenyi Chen
XPENG Technical Report, Arxiv'26  
[arXiv]
[project page]
|
|
|
XPACE: Joint World and Action Modeling from Heterogeneous Experience
Jiacheng Wei*§,
Jerry Bai*§,
Xiaoyu Yue*,
Zidong Wang*,
Xiaoyang Guo*†,
Cheng Chen,
Fanqi Pu,
Fan Wu,
Zhixu Yue,
Yizhuo Li,
Feng Qiu,
Bo Liu,
Yuying Ge,
Hui Zhou,
Chenyi Chen,
Yixiao Ge‡
XPENG Technical Report, Arxiv'26  
[arXiv]
[project page]
For XPACE: * core contribution; § equal contribution; † project lead; ‡ supervision.
|
|
|
|
|
|
EPS3D: End-to-End Feed-Forward 3D Panoptic Segmentation
Runsong Zhu,
Jiaxin Guo,
Xiaoyang Guo†,
Zhengzhe Liu,
Ka-Hei Hui,
Wei Yin,
Kai Chen,
Wei Chen,
Weiqiang Ren,
Yunhui Liu,
Pheng-Ann Heng,
Chi-Wing Fu
International Conference on Machine Learning, ICML'26  
[arXiv]
[code]
|
|
|
Anchor3R: Streaming 3D Reconstruction with Transient Anchors for Long-Horizon Visual Mapping
Peilin Tao,
Chong Cheng,
Yuansen Du,
Caiwei Song,
Zhengqing Chen,
Xiaoyang Guo,
Wei Yin,
Weiqiang Ren,
Qian Zhang,
Hainan Cui,
Shuhan Shen
Preprint, Arxiv'26  
[arXiv]
|
|
|
HorizonStream: Long-Horizon Attention for Streaming 3D Reconstruction
Chong Cheng,
Peilin Tao,
Nanjie Yao,
Guanzhi Ding,
Xianda Chen,
Yuansen Du,
Xiaoyang Guo,
Wei Yin,
Weiqiang Ren,
Qian Zhang,
Zhengqing Chen,
Hao Wang
Preprint, Arxiv'26  
[arXiv]
|
|
|
HorizonDrive: Self-Corrective Autoregressive World Model for Long-horizon Driving Simulation
Conglang Zhang,
Yifan Zhan,
Qingjie Wang,
Zhanpeng Ouyang,
Yu Li,
Zihao Yang,
Xiaoyang Guo,
Weiqiang Ren,
Qian Zhang,
Zhen Dong,
Yinqiang Zheng,
Wei Yin,
Zhengqing Chen
Preprint, Arxiv'26  
[arXiv]
|
|
|
Composing Driving Worlds through Disentangled Control for Adversarial Scenario Generation
Yifan Zhan,
Zhengqing Chen,
Qingjie Wang,
Zhuo He,
Muyao Niu,
Xiaoyang Guo,
Wei Yin,
Weiqiang Ren,
Qian Zhang,
Yinqiang Zheng
European Conference on Computer Vision, ECCV'26  
[arXiv]
[project page]
[code]
|
|
|
Scal3R: Scalable Test-Time Training for Large-Scale 3D Reconstruction
Tao Xie,
Peishan Yang,
Yudong Jin,
Yingfeng Cai,
Wei Yin,
Weiqiang Ren,
Qian Zhang,
Wei Hua,
Sida Peng,
Xiaoyang Guo†,
Xiaowei Zhou†
Conference on Computer Vision and Pattern Recognition, CVPR'26  
(Highlight)
[arXiv]
[project page]
[code]
|
|
|
|
|
|
LiteVGGT: Boosting Vanilla VGGT via Geometry-Aware Cached Token Merging
Zhijian Shu,
Cheng Lin,
Tao Xie,
Wei Yin,
Ben Li,
Zhiyuan Pu,
Weize Li,
Yao Yao,
Xun Cao,
Xiaoyang Guo,
Xiao-Xiao Long
Conference on Computer Vision and Pattern Recognition, CVPR'26  
[arXiv]
[project page]
|
|
|
|
|
|
|
|
|
ComDrive: Comfort-Oriented End-to-End Autonomous Driving
Junming Wang,
Xingyu Zhang,
Zebin Xing,
Songen Gu,
Xiaoyang Guo,
Yang Hu,
Ziying Song,
Qian Zhang,
Xiaoxiao Long,
Wei Yin
International Conference on Intelligent Robots and Systems, IROS'25  
[arXiv]
[project page]
|
|
|
|
|
|
RAD: Training an End-to-End Driving Policy via Large-Scale 3DGS-Based Reinforcement Learning
Hao Gao,
Shaoyu Chen,
Bo Jiang,
Bencheng Liao,
Yiang Shi,
Xiaoyang Guo,
Yuechuan Pu,
Haoran Yin,
Xiangyu Li,
Xinbang Zhang,
Ying Zhang,
Wenyu Liu,
Qian Zhang,
Xinggang Wang
Advances in Neural Information Processing Systems, NeurIPS'25  
[arXiv]
[project page]
|
|
|
Epona: Autoregressive Diffusion World Model for Autonomous Driving
Kaiwen Zhang,
Zhenyu Tang,
Xiaotao Hu,
Xingang Pan,
Xiaoyang Guo,
Yuan Liu,
Jingwei Huang,
Li Yuan,
Qian Zhang,
Xiaoxiao Long,
Xun Cao,
Wei Yin
International Conference on Computer Vision, ICCV'25  
[arXiv]
[project page]
[code]
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|