avatar

Xin Zhou

Second-Year Ph.D. Student
Huazhong University of Science & Technology (HUST)



About Me

I am a second-year Ph.D. student at Huazhong University of Science and Technology, under the supervision of Prof. Xiang Bai. My research interests mainly lie in 3D vision, especially in 3D understanding and generative models.

Research interests: foundation models, vision-language-action (VLA) models, world models, and world action models (WAM) for autonomous driving and embodied AI.

News

Scroll for earlier updates
  • [Sep. 2026] Our work Qwen-Drive-1.0, a vision-language foundation model for autonomous driving, is released.
  • [Sep. 2026] HERMES++ is accepted to TPAMI.
  • [Jul. 2026] Selected as an Outstanding Reviewer for ECCV 2026.
  • [Jun. 2026] One paper is accepted to ECCV 2026.
  • [May. 2026] Recognized as a Silver Reviewer for ICML 2026.
  • [Apr. 2026] Received an Excellence in Reviewing Award from Pattern Recognition.
  • [Feb. 2026] Two papers are accepted to CVPR 2026.
  • [Oct. 2025] Selected as the Top Reviewer at NeurIPS 2025.
  • [Sep. 2025] One paper is accepted to NeurIPS 2025.
  • [Jul. 2025] One paper is accepted to TPAMI 2025.
  • [Jul. 2025] We release a training-free diffusion acceleration framework, now natively integrated into ComfyUI.
  • [Jun. 2025] Finally achieve the bachelor's degree.
  • [Jun. 2025] HERMES is accepted to ICCV 2025.
  • [Apr. 2025] We win the 2-nd place in the SoccerNet Monocular Depth Estimation Challenges (CVPR 2025).
  • [Feb. 2025] One paper is accepted to CVPR 2025.
  • [Sep. 2024] Two papers are accepted to NeurIPS 2024.
  • [Sep. 2024] We win the 1-st place in the FishNet Classification Challenge (ECCV 2024).
  • [Apr. 2024] Invited to give a talk at FALML 2024.
  • [Feb. 2024] One paper is accepted to CVPR 2024!
  • [Jan. 2024] I will keep updating a collection of World Models 2,254 — check it out and give a star🌟!
  • [Aug. 2023] One paper is accepted to PRCV 2023.

Selected Publications

Full list

* Equal contribution

Vision(-Language)(-Action)

Vision-language(-action) foundation model pretraining for autonomous driving and embodied AI, plus unified generation and understanding.

  1. Qwen-Drive-1.0 overview arXiv'26
    Xin Zhou, Zongchuang Zhao, Zhibo Yang, Mingsheng Li, Humen Zhong, Shuai Bai, Du Chu, Ruizhe Chen, Zhaohai Li, Jun Tang, Qiuyue Wang, Mingkun Yang, Jiazhao Zhang, Dayiheng Liu, Dingkang Liang, Xiang Bai
    A vision-language foundation model for autonomous driving that unifies 3D perception, driving VQA, and motion planning.
  2. HERMES++ framework for scene generation and understanding TPAMI'26/ICCV'25
    Xin Zhou, Dingkang Liang, Sifan Tu, Xiwu Chen, Yikang Ding, Dingyuan Zhang, Feiyang Tan, Hengshuang Zhao, Xiang Bai
    A unified driving world model enabling joint multi-view understanding and future lidar generation.

Generative World Model

Driving and general-purpose world models, world action models, and efficient video generation.

  1. EasyCache qualitative comparison against video diffusion caching baselines arXiv'25
    Xin Zhou, Dingkang Liang, Kaijin Chen, Tianrui Feng, Xiwu Chen, Hongkai Lin, Yikang Ding, Feiyang Tan, Hengshuang Zhao, Xiang Bai
    A training-free runtime-adaptive caching scheme that accelerates video diffusion; officially integrated into ComfyUI, plug-and-play for the Wan series, MiniMax H3, and more.

3D Perception

Point cloud and 3D scene understanding.

  1. TPAMI'25
    Dingkang Liang*, Tianrui Feng*, Xin Zhou*, Yumeng Zhang, Zhikang Zou, Xiang Bai
    A spectral domain perspective for point PEFT, very strong performance.
  2. CVPR'24
    Xin Zhou*, Dingkang Liang*, Wei Xu, Xingkui Zhu, Yihan Xu, Zhikang Zou, Xiang Bai
    Introduce Dynamic Adapter with prompt tuning in point PEFT.
  3. NeurIPS'24
    Dingkang Liang*, Xin Zhou*, Wei Xu, Xingkui Zhu, Zhikang Zou, Xiaoqing Ye, Xiao Tan, Xiang Bai
    A simple state space model tailored for point cloud analysis.

Selected Experience

Participated Projects

Academic Service