Tingting Liao

Tingting Liao

廖婷婷

PhD Candidate, MBZUAI

tingting.liao@mbzuai.ac.ae

Scroll

Research

I work on world models for embodied AI — real-time, action-conditioned video generation, and using generated worlds as data engines and evaluators for robot policies. I came to this from generative video and 3D, building controllable models of humans, objects and scenes, which is where my interest in physical consistency and controllability comes from.

World Models  ·  Video Generation  ·  Embodied AI

Education

PhD·MBZUAI2023 – PresentAdvisor: Hao Li

MSc·CASIA2020 – 2023

BEng·Wuhan Polytechnic University2015 – 2019

News

  • MMGR accepted to EMNLP 2026
  • Steering Video Diffusion Transformers with Massive Activations released on arXiv
  • Started a research internship at the Institute of Foundation Models, MBZUAI, on real-time interactive world models
  • Character Mixing for Video Generation released on arXiv
  • Started a research internship at Adobe Research
  • SOAP accepted to SIGGRAPH 2025
  • TADA! and TeCH accepted to 3DV 2024

Experience

UTS

University of Technology Sydney

Visiting Student

2018.12 — 2019.03

Xiaohongshu

Xiaohongshu

Research Intern

2022.02 — 2023.08

Westlake University

Westlake University

Visiting Student

2025.02 — 2025.04

Adobe

Adobe Research

Research Intern

2025.05 — 2025.08

Institute of Foundation Models

IFM, MBZUAI

Research Intern

2026.01 — Present

Publications

Steering Video Diffusion Transformers with Massive Activations

Xianhang Cheng, Yujian Zheng, Zhenyu Xie, Tingting Liao, Hao Li

arXiv 2026

Page Paper Code

Massive activations in video DiTs concentrate on first-frame and latent-boundary tokens. Steering them at inference improves temporal coherence — training-free, with no extra compute.

MMGR: Multi-Modal Generative Reasoning

Zefan Cai, Haoyi Qiu, Tianyi Ma, Haozhe Zhao, Gengze Zhou, Tingting Liao, Xinyan Velocity Yu, Kung-Hsiang Huang, Ke Wan, Shawn Lin, Parisa Kordjamshidi, Minjia Zhang, Wen Xiao, Jiuxiang Gu, Nanyun Peng, Junjie Hu

EMNLP 2026

Page Paper Code

A benchmark that asks generative models to reason rather than render: abstract puzzles, embodied navigation and physical commonsense, scored across five reasoning abilities.

Character Mixing for Video Generation

Tingting Liao, Chongjian Ge, Guangyi Liu, Hao Li, Yi Zhou

arXiv 2025

Page Paper Code

Places characters from different sources into a single video, keeping each one's identity and style intact while they interact.

SOAP: Style-Omniscient Animatable Portraits

Tingting Liao, Yujian Zheng, Adilbek Karmanov, Liwen Hu, Leyang Jin, Yuliang Xiu, Hao Li

SIGGRAPH 2025

Page Paper Video Code

Turns one portrait — photo, painting or cartoon — into a rigged, animatable 3D head, whatever the style.

TADA! Text to Animatable Digital Avatars

Tingting Liao, Hongwei Yi, Yuliang Xiu, Jiaxiang Tang, Yangyi Huang, Justus Thies, Michael J. Black

3DV 2024

#3 Most Influential 3DV 2024 Paper

Page Paper Video Code

Generates animatable 3D avatars from text alone, with high-quality geometry and texture.

TeCH: Text-guided Reconstruction of Lifelike Clothed Humans

Yangyi Huang, Hongwei Yi, Yuliang Xiu, Tingting Liao, Jiaxiang Tang, Deng Cai, Justus Thies

3DV 2024

#6 Most Influential 3DV 2024 Paper

Page Paper Code

Reconstructs a fully clothed human from a single image, using text guidance to hallucinate the regions the camera never saw.

High-Fidelity Clothed Avatar Reconstruction from a Single Image

Tingting Liao, Xiaomei Zhang, Yuliang Xiu, Hongwei Yi, Xudong Liu, Guo-Jun Qi, Yong Zhang, Xuan Wang, Xiangyu Zhu, Zhen Lei

CVPR 2023

Page Paper Code

Recovers a detailed clothed avatar from one image, combining an implicit surface with parametric body priors.

Open Source

Dream2DGS

Image-to-3D with 2D Gaussian splatting — surfel-based reconstruction that keeps geometry crisp where volumetric splatting goes soft.

Code

MirrorGS

Re-implementation of mirror-aware Gaussian splatting, so reflective surfaces reconstruct as mirrors instead of as holes in the scene.

Code

DreamScene360

Re-implementation of text to 360° 3D scenes — a full panoramic environment generated from a single prompt.

Code