Hi there! I'm a third-year undergraduate student at Peking University, majoring in Computer Science.
Currently, I am a visiting student at BAIR at UC Berkeley, advised by Prof. Trevor Darrell.
Prior to this, I was a research intern at the HMI Lab at Peking University, advised by Prof. Shanghang Zhang.
My research interests lie in the application of multimodal large models, specifically VLA models, in robot manipulation.
[2026.06] I am awarded the SenseTime Scholarship 2026! (30 recipients nationwide)
[2026.04] One paper (LaST0) is accepted to ICML 2026 as a spotlight (top 2.2%)!
[2026.02] One paper (ManualVLA) is accepted to CVPR 2026!
[2026.01] One paper (MLA) is accepted to ICRA 2026!
[2026.01] One paper (HybridVLA) is accepted to ICLR 2026!
[2025.09] Two papers (Fast-in-Slow, AC-DiT) are accepted to NeurIPS 2025!
[2025.08] One paper (3DS-VLA) is accepted to CoRL 2025!
Selected Publications
I'm interested in Computer Vision, Robot Learning and Embodied Large Multimodal Models. My research focuses on how to enhance large multimodal models to better reason about the physical world and develop effective task planning. Some papers are highlighted.
T-Rex introduces a large-scale tactile dataset of 100 hours and a variable-rate Mix-of-Transformer (MoT) architecture with a temporal tactile VQ-VAE encoder, achieving ~30% higher success rates across 12 real-world tasks involving force control and deformable object handling.
A VLA model that enables efficient reasoning before acting through a Latent Spatio-Temporal Chain-of-Thought (CoT), capturing fine-grained physical and robotic dynamics that are often difficult to verbalize.
A unified VLA framework built upon a Mixture-of-Transformers (MoT) architecture, enabling coherent collaboration between multimodal manual generation and action execution.
A multisensory language-action (MLA) model that collaboratively perceives heterogeneous sensory modalities and predicts future multisensory objectives to facilitate physical world modeling.
Unlike previous dual-system VLA methods that attach a separate policy head as System 1, FiS-VLA repurposes the final transformer blocks of an intact VLM as System 1, while retaining the full model for System 2 reasoning.
HybridVLA innovatively integrates diffusion and autoregressive action prediction within a single LLM, fully leveraging the continuity and probabilistic nature of diffusion alongside the reasoning capabilities of autoregressive modeling.
Education
University of California, Berkeley Visiting Student, Berkeley Global Access (BGA) Program
2026.01 - Present
Peking University B.S. in Computer Science, Yuanpei College
2023.08 - Present
Research Experience
UC Berkeley Research Intern, Berkeley AI Research
2026.01 - Present
Research on Embodied AI and Robot Manipulation
Research Advisor: Prof. Trevor Darrell
Simplexity Robotics Research Intern
2025.09 - 2026.01
Research on Embodied AI and Robot Manipulation
Beijing Innovation Center of Humanoid Robotics Research Intern
2025.08 - 2026.01
Research on Embodied AI and Robot Manipulation
AI2Robotics Research Intern, X-Lab
2025.06 - 2026.01
Focused on Vision-Language-Action (VLA) models
Peking University Research Intern, HMI (Human Machine Intelligence) Lab
2024.07 - Present