Events
Computer Science Department ColloquiumFoundations of Generative AI: Data Dependence in Diffusion Models |
|
||
Wednesday, March 04, 2026, 10:30am - 12:00pm |
|||
Speaker: Chris Scarvelis
Bio
Chris Scarvelis is a final-year PhD student in the Geometric Data Processing Group at MIT, where he works on the theoretical foundations of generative AI. His research focuses on understanding why modern generative models produce novel content rather than memorizing their training data, and how a model’s behavior is shaped by its training data. His work has been supported by an MIT–Google Future Research Cohort Fellowship, an Exponent Fellowship, a Siebel Scholarship, and an NSERC PGS-D Scholarship. Beyond his core research in generative modeling, Chris has worked on graph neural networks at Twitter, 3D deep learning at Backflip AI, and model merging methods for large language models at the MIT-IBM Watson AI Lab.
Location : CoRE 301
Committee:
Event Type: Computer Science Department Colloquium
Abstract: Diffusion models are fundamental tools in generative AI, powering systems that synthesize data ranging from images to video to molecules. However, their empirical success has been accompanied by striking theoretical paradoxes. For example, training a diffusion model to optimality yields a model that can only generate training data, and diffusion models succeed in practice precisely because they fail to reach this trivial optimum. This gap between theory and practice highlights how little we know about the relationship between diffusion models’ real-world behavior and their training data.This talk will describe my recent efforts toward closing this gap. In the first part of the talk, I will explore how neural networks’ smoothness bias enables neural diffusion models to generate new samples, and how this perspective leads to a training-free generative model obtained by simulating this bias directly. I will then introduce a regularizer that captures diffusion models’ bias toward learning locally low-rank functions and show how to exploit the compositional structure of neural networks to implement this regularizer efficiently. Zooming in on the impact of the training data on a particular trained model’s behavior, I will finally describe a framework for differentiating diffusion models with respect to their target distributions. This sensitivity analysis reveals how a diffusion model’s score function and generated samples respond to infinitesimal perturbations of its training set.Together, these results build toward a predictive theory of generative AI that explains the real-world behavior of generative models via neural architectures, training procedures, and the geometry of their training data. Beyond theory, this framework has practical implications for data attribution, targeted data acquisition, and the design of more interpretable and controllable generative models.
Organization:
Contact Professor Lirong Xia
Join Zoom Meeting
https://rutgers.zoom.us/j/2014444359?pwd=WW9ybFNCNVFrUWlycHowSHdNZjhzUT09
Meeting ID: 201 444 4359
Password: 550978