CS Events
PhD DefenseEnhancing Visual Understanding with Large Foundational Models |
|
||
Wednesday, April 30, 2025, 03:00pm - 05:00pm |
|||
Speaker: Shiyu Zhao
Bio
Location : CoRE 301
Committee:
Professor Dimitris N. Metaxas
Professor Konstantinos Michmizos
Professor Ruixiang Tang
Dr. Han Zhang (external)
Event Type: PhD Defense
Abstract: Large foundation models, such as vision and language models (VLMs), large language models (LLMs), and multimodal large language models (MLLMs), have demonstrated remarkable potential across a wide range of tasks, pushing us closer to achieving artificial general intelligence. However, their large model sizes pose significant efficiency challenges, particularly in vision and language tasks that require processing high-resolution images. Moreover, directly applying these models to domain-specific tasks, such as open-vocabulary object detection (OVOD), remains a non-trivial problem due to domain-specific constraints and data limitations. During my PhD program, I present a series of studies aimed at addressing efficiency and domain-specific challenges. To tackle the efficiency issue, I propose a search-based algorithm that optimizes the removal of redundant and unnecessary computations related to vision tokens. My approach accelerates MLLMs by 2× without a performance drop. Compared to existing methods, it achieves a superior efficiency-performance trade-off when the computational budgets are constrained. For domain-specific tasks, I explore strategies to leverage large foundation models to enhance domain-specific models, with a focus on OVOD. Instead of directly applying foundation models, my research investigates how these models can be utilized to address data limitations in OVOD. Specifically, I identify critical gaps in the training data that are either missing or challenging to collect. By generating synthetic data through large foundation models to fill these gaps, I demonstrate significant performance improvements in OVOD models.In a nutshell, my research improves the efficiency of large foundation models and enhances their applicability to domain-specific tasks, ultimately making them more practical and impactful in real-world scenarios.
Organization:
Contact Dimitris N. Metaxas
https://rutgers.zoom.us/j/93986623264?pwd=OJZMYEc3fYsG3OL6tLs0beRQXXlV38.1&jst=2
Subscribe to RSS Feed