Real2SAM2Real: Generative 3D Caches as Complementary Context for Video Diffusion
Manuscript
Control the camera and the objects in video generation with a generative 3D cache
On the job market · full-time from 2027
PhD Student · UMD · expected 2027
CS PhD student at the University of Maryland, advised by Christopher Metzler.
I work on making image and video generation physically controllable. Diffusion models can produce convincing pictures, but they give you few reliable handles on how a picture was formed. My work adds those handles (where the light comes from, what is in focus, what is reflected in glass) and uses 3D structure and other sensors, such as event cameras, to keep the result consistent with a real scene. I have also worked on restoring images captured in hard conditions and on perceptual image quality.
Previously at X-Pixel (2020–2022) with Chao Dong and Jinjin Gu; BEng from The Chinese University of Hong Kong, Shenzhen. Always happy to talk research or early-career questions.
Two figures from my papers — move your cursor across them to move the light, and to pull focus.
ICCV 2025
Put a shadow knob on a text-to-image model — direction, softness and shape, with no retraining. Move your cursor across to walk the light around her face.
ECCV 2026
Render geometrically consistent bokeh on scenes that defeat other methods — here, a toy behind a transparent lid. Move your cursor across to pull focus through it.
Our work on vision–language robustness under adverse imaging was accepted to CoLM 2026. Congratulations to Tianfu and Mingyang.
Large-Scale Light Field Synthesis from Videos was accepted to ECCV 2026.
Back at the Adobe NextCam team for a second summer internship.
Parametric Shadow Control was accepted to ICCV 2025.
Named an Outstanding Reviewer for CVPR 2025.
Joined the Adobe NextCam team (led by Marc Levoy) as a summer intern, mentored by Shumian Xin and Zhoutong Zhang.
Two papers accepted to CVPR 2025. Congratulations to Tianfu, Mingyang, and Jingxi.
Summer 2026
Ph.D. Research Intern, NextCam team (Marc Levoy).
Summer 2025
Ph.D. Research Intern, NextCam team (Marc Levoy), supervised by Shumian Xin and Zhoutong Zhang.
Bokeh editing project; outcome: ECCV 2026. Patent filing in process.
Summer 2024
Ph.D. Research Intern, supervised by Guan-Ming Su.
Portrait lighting control for text-to-image diffusion models without light stage data. Outcome: ICCV 2025, one patent.
2020 – 2022
Undergraduate Research Assistant, supervised by Chao Dong.
Efficient and controllable image restoration; image quality and aesthetic assessment. Outcome: two ECCV papers, two CVPR workshops organized, one patent.
* denotes equal contribution.
Manuscript
Control the camera and the objects in video generation with a generative 3D cache
ECCV 2026
Render geometrically consistent bokeh on scenes that defeat other methods — transparent surfaces, thin structures, cluttered depth
ICCV 2025
Add a shadow knob to AI portraits, without heavy compute
CVPR 2025
Separate reflection from transmission with flash cues and a diffusion prior
CVPR 2025
Turn event streams into smooth video with a video diffusion model
NeurIPS 2024
Low-pass the temporal dimension to strip atmospheric turbulence out of video
ECCV 2024
Separate reflections in 3D using flash-induced cues

CVPR 2024
Engineer the PSF so event cameras encode depth for free
ICCV 2023
Remove snow from video consistently across frames, learned from real snowfall rather than synthetic

NeurIPS'23
Assess 360° image quality the way a viewer actually explores the sphere

ECCV 2020
Score image quality on the artifacts that modern restoration models actually produce
PDFDatasetNTIRE'21 ChallengeNTIRE'22 Full-RefNTIRE'22 No-Ref
AIM Challenge @ ECCV 2022
NTIRE Challenge @ CVPR 2022
Presides over the Maryland headquarters. Contributes to every paper by lying on the keyboard. Has strong feelings about not being in the acknowledgements.





