Seung Hyun Lee · Ph.D. Candidate

Learning structure for visual and embodied intelligence.

I develop structured models that connect perception, generative modeling, and physical interaction.

I am a Ph.D. candidate in Computer Science and Engineering at the University of Michigan, advised by Prof. Stella X. Yu. My research focuses on foundation models for visual and embodied intelligence, spanning structure-aware perception, generative modeling, and robot learning.

My research experience includes Google DeepMind, where I worked on embodied intelligence and robot learning, and Google Research, where I developed methods for controllable image generation and image cropping. Previously, I received my M.S. in Artificial Intelligence from Korea University.

PRISM was accepted to CoRL 2026.View project ↗

Selected Work

CoRL 2026 · Accepted

PRISM: Polynomial Representations for Interaction-Structured Motor Control

PRISM makes polynomial interactions among observable physical variables explicit and learnable, improving locomotion and contact-rich manipulation without adding sensors.

Seung Hyun Lee, Stella X. Yu

SHED uses a learned segment hierarchy to produce structurally coherent dense predictions
arXiv 2026

SHED Light on Segmentation for Dense Prediction

SHED learns a bidirectional hierarchy of segment tokens without segmentation supervision, producing sharper depth boundaries, coherent scene structure, and stronger synthetic-to-real generalization.

Seung Hyun Lee, Sangwoo Mo, Stella X. Yu

Selected Publications

CVPR 2025

Cropper: Vision-Language Model for Image Cropping through In-Context Learning

Seung Hyun Lee*, Jijun Jiang*, Yiran Xu*, Zhuofang Li*, Junjie Ke, Yinxiao Li, Junfeng He, Steven Hickson, Katie Datsenko, Sangpil Kim, Ming-Hsuan Yang, Irfan Essa, Feng Yang

ECCV 2024
Oral

Parrot: Pareto-optimal Multi-Reward Reinforcement Learning Framework for Text-to-Image Generation

Seung Hyun Lee, Yinxiao Li, Junjie Ke, Innfarn Yoo, Han Zhang, Jiahui Yu, Qifei Wang, Fei Deng, Glenn Entis, Junfeng He, Gang Li, Sangpil Kim, Irfan Essa, Feng Yang

CVPR 2022

Sound-Guided Semantic Image Manipulation

Seung Hyun Lee, Wonseok Roh, Wonmin Byeon, Sang Ho Yoon, Chanyoung Kim, Jinkyu Kim*, Sangpil Kim*

ECCV 2022

Sound-Guided Semantic Video Generation

Seung Hyun Lee, Gyeongrok Oh, Wonmin Byeon, Chanyoung Kim, Won Jeong Ryoo, Sang Ho Yoon, Hyunjun Cho, Jihyun Bae, Jinkyu Kim*, Sangpil Kim*

Neural Networks
2024

Robust Sound-Guided Image Manipulation

Seung Hyun Lee*, Hyung-gun Chi*, Gyeongrok Oh, Wonmin Byeon, Sang Ho Yoon, Hyunje Park, Wonjun Cho, Jinkyu Kim*, Sangpil Kim*

CVM 2024

Audio-Guided Implicit Neural Representation for Local Image Stylization

Seung Hyun Lee*, Sieun Kim*, Wonmin Byeon, Gyeongrok Oh, Sumin In, Hyeongcheol Park, Sang Ho Yoon, Sung-Hee Hong, Jinkyu Kim, Sangpil Kim

arXiv

Soundini: Sound-Guided Diffusion for Natural Video Editing

Seung Hyun Lee, Sieun Kim, Innfarn Yoo, Feng Yang, Donghyeon Cho, Youngseo Kim, Huiwen Chang, Jinkyu Kim*, Sangpil Kim*

Teaching

  • Fall 2026

    CSE 598: 3D Visual Representation

    Lecturer

  • Fall 2026

    CSE 592: Foundations of Artificial Intelligence

    Lecturer & Graduate Student Instructor

    Course website
  • Winter 2026

    EECS 542 · Guest Lecture: Depth Foundation Models

    Guest Lecturer

  • Winter 2026

    EECS 442: Computer Vision

    Graduate Student Instructor

  • Spring 2023

    Machine Learning, Korea University

    Teaching Assistant