计算机学院系列讲座菁英论坛第51期——Towards Spatial Intelligence in Vision Language Models: Benchmarking, Analysis and Improvement
报告题目(Title):Towards Spatial Intelligence in Vision Language Models: Benchmarking, Analysis and Improvement
时间 (Date & Time):2026.07.07 10:00 am – 11:00 am
地点 (Location):燕园大厦 809, room 809 Yan Yuan Building
主讲人 (Speaker):Wei Gao (University of Pittsburgh)
邀请人 (Host) :Chenren Xu
报告摘要 (Abstract):
Spatial intelligence is the cognitive ability to mentally perceive and manipulate objects in the 3D space. In recent embodied AI and robotic applications, spatial intelligence is the key foundation for AI to autonomously perceive the surrounding physical world and comprehend the spatial relationship of objects as humans do, hence allowing correct interpretation of spatial contexts and guide device actions. Although today's Vision Language Models (VLMs) perform well in traditional visual tasks such as object detection and recognition, they perform poorly in spatial intelligence tasks that involve visual spatial perception and reasoning. In this talk, I will begin by introducing the fundamental difference between spatial intelligence and traditional language intelligence, which result in the poor performance of current VLMs in spatial intelligence tasks. Then, I will present our recent works on studying and advancing spatial intelligence in today's VLMs. First, I will present InfiniBench, a fully customizable framework that allows multi-view spatial benchmark synthesis with parameterized control. Second, using a broad variety of benchmarks generated by InfiniBench, I will present our results of VLM reasoning path and latent state analysis that unveil the VLMs' failure modes in spatial intelligence tasks. Third, I will show how such analysis results motivate us to investigate into VLMs' latent embedding space, to identify, extract and utilize subspaces that only contain spatially relevant aspects to improve VLMs' performance in spatial intelligence.
主讲人简介(Bio):

Wei Gao is currently an Associate Professor in the Department of Electrical and Computer Engineering, University of Pittsburgh. His research is positioned at the intersection among computer systems, AI, and computer vision, focusing on the design and deployment of multimodal generative AI, world action and reasoning models, and on-device AI systems, architectures and algorithms. He also has strong interests in unveiling the analytical principles underneath AI deployment problems, and designing AI systems that can ensure accuracy, adaptability and generalizability in applications based on these principles. He has published more than 100 research papers at top AI, computer vision and system conference venues, including ICLR, AAAI, ICCV, CVPR, ECCV, ASPLOS, MobiCom, MobiSys, SenSys, etc, and received multiple best paper awards or nominations.

欢迎关注计算机学院微信公众号,了解更多讲座信息!
北京大学计算机学院
