An-Chieh Cheng outdoors

An-Chieh Cheng

鄭安傑

PhD student, UC San Diego

I am a PhD student at the University of California, San Diego, advised by Prof. Xiaolong Wang. During my PhD studies, I interned at NVIDIA and Adobe, and my research has been supported by Qualcomm Innovation Fellowship. Prior to my PhD, I earned my Master’s and Bachelor’s degrees in computer science from National Tsing Hua University.

Research directions

I'm interested in building multimodal foundation models that understand space, act intelligently, and self-evolve through real-world experience.

News

Gave talks on “Grounding 3D Reasoning and Long-Horizon Actions in VLMs and VLAs” at the CVPR 2026 Workshop on Efficient Deep Learning for Computer Vision () and Foundation Models Meet Embodied Agents in Denver. Thank you for having me!
Excited to share Cosmos 3, NVIDIA's omnimodal world foundation model powering the next generation of Physical AI — grateful to have contributed to its spatial reasoning capabilities.
GR3D is accepted by CVPR 2026.
Earlier updates
Released code for SR-3D, together with SR-3D-Bench.
Two papers (SR3D and OmniVinci) accepted by ICLR 2026.
We’ve open-sourced the NaVILA framework. Including the VLA, locomotion policy, and the navigation benchmark.
SpatialRGPT was demoed at GTC 2025 as a part of Agentic AI for Physical Operations!

Publications & preprints

Full list ↗

Long-Horizon Manipulation via Trace-Conditioned VLA Planning

Isabella Liu, An-Chieh Cheng, Rui Yan, Geng Chen, Ri-Zhao Qiu, Xueyan Zou, Sha Yi, Hongxu Yin, Xiaolong Wang, Sifei Liu
CoRL, 2026

Long-horizon manipulation via a task-management VLM with visual trace conditioning.

Grounded 3D-Aware Spatial Vision-Language Modeling

An-Chieh Cheng, Yang Fu, Yatai Ji, Ligeng Zhu, Guanqi Zhan, Zhuoyang Zhang, Zhaojing Yang, Song Han, Yao Lu, Pavlo Molchanov, Vidya Nariyambut Murali, Jan Kautz, Xiaolong Wang, Hongxu Yin, Sifei Liu
CVPR, 2026

Unified Spatial Reasoning & 3D Grounding VLMs with visual CoT (Thinking with Regions).

OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLM

Hanrong Ye*, C.-H. Huck Yang, Arushi Goel, Wei Huang, Ligeng Zhu, Yuanhang Su, Sean Lin, An-Chieh Cheng, Zhen Wan, Jinchuan Tian, Yuming Lou, Dong Yang, et al.
ICLR, 2026

NVIDIA's state-of-the-art 9B Omni-Modal LLMs.

3D Aware Region Prompted Vision Language Model

An-Chieh Cheng, Yang Fu, Yukang Chen, Zhijian Liu, Xiaolong Li, Subhashree Radhakrishnan, Song Han, Yao Lu, Jan Kautz, Pavlo Molchanov, Hongxu Yin✝︎, Xiaolong Wang✝︎, Sifei Liu✝︎
ICLR, 2026

Region-level spatial reasoning for both single-view and multi-view inputs.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Ruihan Yang*, Qinxi Yu*, Yecheng Wu, Rui Yan, Borui Li, An-Chieh Cheng, Xueyan Zou, Yunhao Fang, Xuxin Cheng, Ri-Zhao Qiu, Hongxu Yin, Sifei Liu, Song Han, Yao Lu, Xiaolong Wang
IROS, 2026

Robust dexterous manipulation generalist model utilizing diverse egocentric human manipulation videos.

NVILA teaser

NVILA: Efficient Frontier Visual Language Models

Zhijian Liu et al.
CVPR, 2025

Efficient frontier VLM models with efficient training and inference.

SpatialRGPT: Grounded Spatial Reasoning in Vision-Language Models

An-Chieh Cheng, Hongxu Yin, Yang Fu, Qiushan Guo, Ruihan Yang, Jan Kautz, Xiaolong Wang, Sifei Liu
NeurIPS, 2024

A powerful region-level VLM adept at 3D spatial reasoning.
✨ Demoed at GTC 2025 as a part of Agentic AI for Physical Operations!

TUVF: Learning Generalizable Texture UV Radiance Fields

An-Chieh Cheng, Xueting Li, Sifei Liu✝︎, Xiaolong Wang✝︎
ICLR, 2024

Learning generalizable texture UV radiance fields for shapes.

Autoregressive 3D Shape Generation teaser

Autoregressive 3D Shape Generation via Canonical Mapping

An-Chieh Cheng*, Xueting Li*, Sifei Liu, Min Sun, Ming-Hsuan Yang
ECCV, 2022

We decompose the point cloud into meaningful shape sequences, then we encode these sequences through a transformer for generation.

Canonical Point Autoencoder teaser

Learning 3D Dense Correspondence via Canonical Point Autoencoder

An-Chieh Cheng, Xueting Li, Min Sun, Ming-Hsuan Yang, Sifei Liu
NeurIPS, 2021

Unsupervised learning of dense 3D correspondence.

Technical reports

Vesta: A Generalist Embodied Reasoning Model

Johan Bjorck*, Zhiqi Li*, Yunze Man*, Jing Wang*, An-Chieh Cheng, Sifei Liu, Shihao Wang, Zhiding Yu, Abhishek Badki, Stan Birchfield, Valts Blukis, et al.
NVIDIA Technical Report, 2026

A state-of-the-art Robot System 2 embodied-reasoning VLM that unifies localization, navigation, memory, reasoning, tool use, and long-horizon planning.

Cosmos 3: Omnimodal World Models for Physical AI

Nvidia et. al.
NVIDIA Technical Report, 2026

NVIDIA's omnimodal world foundation model that unifies understanding, generation, simulation, and action across text, image, video, audio, and robot actions for Physical AI.