Now I work at Alibaba Token Hub (ATH) Group. I received my Ph.D degree at Artificial Intelligence and Biomedical Image Analysis Lab, School of Engineering, Westlake University (Hangzhou, China), affiliated with a joint-supervision program with Zhejiang University, supervised by Prof. Lin Yang. Before that, I received my B.Eng Degree from School of Automation Science and Electrical Engineering, Beihang University (Beijing, China) in June 2021.

Currently, my research is centered on Visual-Language Model for Spatial Intelligence.

πŸ”₯ News

  • 2025.9: Β πŸŽ‰πŸŽ‰ One paper was accepted by NeurIPS 2025!
  • 2024.11: Β πŸŽ‰πŸŽ‰ Awarded as Zhejiang University 2024 National Scholarship!
  • 2024.10: Β πŸŽ‰πŸŽ‰ Awarded as MICCAI Young Scientist Award (Top 5)!
  • 2024.10: Β πŸŽ‰πŸŽ‰ One paper was accepted by NeurIPS 2024!
  • 2024.09: Β πŸŽ‰πŸŽ‰ One MICCAI 2024 paper was invited as oral presentation!
  • 2024.07: Β πŸŽ‰πŸŽ‰ One paper was accepted by European Conference on Computer Vision (ECCV 2024)!
  • 2024.05: Β πŸŽ‰πŸŽ‰ One paper was Early Accepted by International Conference on Medical Image Computing and Computer Assisted Intervention (MICCAI 2024)!.
  • 2023.12: Β πŸŽ‰πŸŽ‰ Awarded as Zhejiang University 2023 Outstanding Graduates!
  • 2023.05: Β πŸŽ‰πŸŽ‰ One paper was Early Accepted by International Conference on Medical Image Computing and Computer Assisted Intervention (MICCAI 2023)!.
  • 2022.12: Β πŸŽ‰πŸŽ‰ Awarded as Zhejiang University 2022 Outstanding Graduates!
  • 2022.01: Β πŸŽ‰πŸŽ‰ One paper was accepted by IEEE International Symposium on Biomedical Imaging (ISBI 2022)!

πŸ“ Publications

NeurIPS 2025
sym

SD-VLM: Spatial Measuring and Understanding with Depth-Encoded Vision-Language Models (NeurIPS 2025)

Pingyi Chen, Yujing Lou, Shen Cao, Jinhui Guo, Lubin Fan, Yue Wu, Lin F. Yang, Lizhuang Ma, Jieping Ye

Paper GitHub Hugging Face

  • Propose Massive Spatial Measuring and Understanding (MSMU) dataset with precise spatial annotations.
  • Introduce a plug-in depth positional encoding method strengthening VLMs’ spatial awareness.
ECCV 2024
sym

WSI-VQA: Interpreting Whole Slide Images by Generative Visual Question Answering European Conference on Computer Vision (ECCV 2024)

Pingyi Chen, Chenglu Zhu, Sunyi Zheng, Honglin Li, Lin Yang

Paper GitHub

  • Establish a scalable pipeline to curate slide-level VQA datasets.
  • Propose a WSI-VQA framework to reframe previous WSI-related tasks into a QA pattern.
MICCAI 2024
sym

WsiCaption: Multiple Instance Generation of Pathology Reports for Gigapixel Whole-Slide Images International Conference on Medical Image Computing and Computer Assisted Intervention (MICCAI 2024 Oral/Best Paper Candidate)

Pingyi Chen, Honglin Li, Chenglu Zhu, Sunyi Zheng, Zhongyi Shui, Lin Yang

Paper GitHub

  • Establish a pathology report dataset with GPT.
  • Propose a multiple instance generation (MI-Gen) framework to generate text based on gigapixel inputs.

πŸŽ– Honors and Awards

  • 2024.11 Zhejiang University 2024 National Scholarship.
  • 2024.10 MICCAI Young Scientist Award.
  • 2023.12 Zhejiang University 2023 Outstanding Graduates.
  • 2022.12 Zhejiang University 2022 Outstanding Graduates.
  • 2022.06 Westlake University Outstanding Cadres of the Graduate Student Union.

πŸ“– Educations

  • 2021.09 - 2026.06, Ph.D at Westlake University & Zhejiang University, Hangzhou, China.
  • 2017.09 - 2021.06, Bachelor of Engineering, Automation, Beihang University, Beijing, China.

πŸ’¬ Talks

πŸ’» Internships

  • 2024.12 - 2025.10, Tongyi Lab, Alibaba, Hangzhou, China.
  • 2021.03 - 2021.08, SenseTimeοΌˆε•†ζ±€οΌ‰Research, Beijing, China.