Zhuoyang Liu | 刘卓洋

Hi there! I'm a third-year undergraduate student at Peking University, majoring in Computer Science. Currently, I am a visiting student at BAIR at UC Berkeley, advised by Prof. Trevor Darrell. Prior to this, I was a research intern at the HMI Lab at Peking University, advised by Prof. Shanghang Zhang. My research interests lie in the application of multimodal large models, specifically VLA models, in robot manipulation.

Email  /  Scholar  /  Twitter  /  Github  /  WeChat

profile photo

News

  • [2026.06] I am awarded the SenseTime Scholarship 2026! (30 recipients nationwide)
  • [2026.04] One paper (LaST0) is accepted to ICML 2026 as a spotlight (top 2.2%)!
  • [2026.02] One paper (ManualVLA) is accepted to CVPR 2026!
  • [2026.01] One paper (MLA) is accepted to ICRA 2026!
  • [2026.01] One paper (HybridVLA) is accepted to ICLR 2026!
  • [2025.09] Two papers (Fast-in-Slow, AC-DiT) are accepted to NeurIPS 2025!
  • [2025.08] One paper (3DS-VLA) is accepted to CoRL 2025!

Selected Publications

I'm interested in Computer Vision, Robot Learning and Embodied Large Multimodal Models. My research focuses on how to enhance large multimodal models to better reason about the physical world and develop effective task planning. Some papers are highlighted.

T-Rex: Tactile-Reactive Dexterous Manipulation
Dantong Niu*, Zhuoyang Liu*, Zekai Wang*, Boning Shao, Zhao-Heng Yin, Anirudh Pai,
Yuvan Sharma, Stefano Saravalle, Ruijie Zheng, Jing Wang, Ryan Punamiya, Mengda Xu,
Yuqi Xie, Yunfan Jiang, Letian Fu, Konstantinos Kallidromitis, Matteo Gioia, Junyi Zhang, Jiaxin Ge, Haiwen Feng, Fabio Galasso, Wei Zhan, David M. Chan, Yutong Bai, Roei Herzig,
Jiahui Lei, Fei-Fei Li, Ken Goldberg, Jitendra Malik, Pieter Abbeel, Yuke Zhu, Danfei Xu,
Jim (Linxi) Fan, Trevor Darrell
arXiv, 2026
project page / arXiv / code / dataset

T-Rex introduces a large-scale tactile dataset of 100 hours and a variable-rate Mix-of-Transformer (MoT) architecture with a temporal tactile VQ-VAE encoder, achieving ~30% higher success rates across 12 real-world tasks involving force control and deformable object handling.

LaST0: Latent Spatio-Temporal Chain-of-Thought for Robotic Vision-Language-Action Model
Zhuoyang Liu*, Jiaming Liu*, Hao Chen*, Jiale Yu, Ziyu Guo, Chengkai Hou,
Chenyang Gu, Xiangju Mi, Renrui Zhang, Kun Wu, Zhengping Che, Jian Tang,
Pheng-Ann Heng, Shanghang Zhang
ICML, 2026 (spotlight, top 2.2%)
project page / arXiv / code

A VLA model that enables efficient reasoning before acting through a Latent Spatio-Temporal Chain-of-Thought (CoT), capturing fine-grained physical and robotic dynamics that are often difficult to verbalize.

ManualVLA: A Unified VLA Model for Chain-of-Thought Manual Generation and Robotic Manipulation
Chenyang Gu*, Jiaming Liu*, Hao Chen*, Runzhong Huang*, Qingpo Wuwu, Zhuoyang Liu,
Xiaoqi Li, Ying Li, Renrui Zhang, Peng Jia, Pheng-Ann Heng, Shanghang Zhang
CVPR, 2026
project page / arXiv / video

A unified VLA framework built upon a Mixture-of-Transformers (MoT) architecture, enabling coherent collaboration between multimodal manual generation and action execution.

DualVLA: Building a Generalizable Embodied Agent via Partial Decoupling of Reasoning and Action
Zhen Fang*, Zhuoyang Liu*, Jiaming Liu, Hao Chen, Yu Zeng, Shiting Huang,
Zehui Chen, Lin Chen, Shanghang Zhang, Feng Zhao
arXiv, 2025
project page / arXiv

DualVLA improves action performance through carefully designed post-training while preserving the reasoning ability.

MLA: A Multisensory Language-Action Model for Multimodal Understanding and Forecasting in Robotic Manipulation
Zhuoyang Liu*, Jiaming Liu*, Jiadong Xu, Nuowei Han, Chenyang Gu, Hao Chen,
Kaichen Zhou, Renrui Zhang, Kai Chin Hsieh, Kun Wu, Zhengping Che, Jian Tang,
Shanghang Zhang
ICRA, 2026
project page / arXiv / code

A multisensory language-action (MLA) model that collaboratively perceives heterogeneous sensory modalities and predicts future multisensory objectives to facilitate physical world modeling.

AC-DiT: Adaptive Coordination Diffusion Transformer for Mobile Manipulation
Sixiang Chen*, Jiaming Liu*, Siyuan Qian*, Han Jiang, Xiaoqi Li, Renrui Zhang,
Zhuoyang Liu, Chenyang Gu, Chengkai Hou, Pengwei Wang, Zhongyuan Wang, Shanghang Zhang
NeurIPS, 2025
project page / arXiv / code

Adaptive Coordination Diffusion Transformer (AC-DiT) enhances mobile base and manipulator coordination for end-to-end mobile manipulation.

Fast-in-Slow: A Dual-System Foundation Model Unifying Fast Manipulation within Slow Reasoning
Hao Chen*, Jiaming Liu*, Chenyang Gu*, Zhuoyang Liu*, Renrui Zhang, Xiaoqi Li,
Xiao He, Yandong Guo, Chi-Wing FU, Shanghang Zhang, Pheng-Ann Heng
NeurIPS, 2025
project page / arXiv / code

Unlike previous dual-system VLA methods that attach a separate policy head as System 1, FiS-VLA repurposes the final transformer blocks of an intact VLM as System 1, while retaining the full model for System 2 reasoning.

HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model
Jiaming Liu*, Hao Chen*, Zhuoyang Liu*, Pengju An*, Renrui Zhang, Chenyang Gu,
Xiaoqi Li, Ziyu Guo, Sixiang Chen, Mengzhen Liu, Chengkai Hou, Mengdi Zhao,
Kaichen Zhou, Pheng-Ann Heng, Shanghang Zhang
ICLR, 2026
project page / arXiv / code

HybridVLA innovatively integrates diffusion and autoregressive action prediction within a single LLM, fully leveraging the continuity and probabilistic nature of diffusion alongside the reasoning capabilities of autoregressive modeling.



Education

UC Berkeley
University of California, Berkeley
Visiting Student, Berkeley Global Access (BGA) Program
2026.01 - Present
Peking University
Peking University
B.S. in Computer Science, Yuanpei College
2023.08 - Present

Research Experience

BAIR
UC Berkeley
Research Intern, Berkeley AI Research
2026.01 - Present

Research on Embodied AI and Robot Manipulation
Research Advisor: Prof. Trevor Darrell

simplexity
Simplexity Robotics
Research Intern
2025.09 - 2026.01

Research on Embodied AI and Robot Manipulation

X-Humanoid
Beijing Innovation Center of Humanoid Robotics
Research Intern
2025.08 - 2026.01

Research on Embodied AI and Robot Manipulation

AI2Robotics
AI2Robotics
Research Intern, X-Lab
2025.06 - 2026.01

Focused on Vision-Language-Action (VLA) models

HMI Lab
Peking University
Research Intern, HMI (Human Machine Intelligence) Lab
2024.07 - Present

Embodied AI
Research Advisor: Prof. Shanghang Zhang

Honors & Awards


2026
SenseTime Scholarship 2026 (only 30 recipients in China)
2025
Academic Rising Star Award Nomination (Undergraduate Program), Peking University
2025
Third Prize in the first round of the RoboTwin Dual-Arm Collaboration Challenge, CVPR 2025
2024
Top 16 in the Mahjong AI Competition Finals, IJCAI 2024
2023
Qin-Jin Scholarship, Peking University
2022
Gold Medalist, 36th Chinese Chemistry Olympiad; Member of the National Training Squad (Top 50 Nationwide)

This homepage is designed based on Jon Barron's website and deployed on Github Pages.