Tuesday, November 12, 2024, 9am

Foundation models have become central to modern machine learning, introducing new capabilities such as zero-shot and few-shot learning. These paradigms offer significant advantages in terms of generalization and robustness. However, as their applications expand, so do the associated risks and limitations, including potential privacy violations and intellectual property concerns.

This thesis proposal seeks to analyze foundation models through the lens of information theory, focusing on three interconnected areas: robustness, privacy, and copyright protection. First, we investigate how contrastive vision-language models respond to distribution shifts, where training and test data differ significantly, proposing information-theoretic measures to quantify and enhance model robustness. Second, we examine how existing unlearning algorithms fail to remove private information from models, and develop efficient auditing tools for machine unlearning. Third, we explore the memorization issues in large language models, advocating for a compression-based adversarial prompt method to measure memorization, which is essential in identifying copyright infringement.

For future work, we first extend our afore-mentioned memorization measurement to continuous soft token space. Previously, this measurement requires a computation in the discrete space, which is not computationally efficient. We aim to achieve a more efficient algorithm by searching in the continuous space, and develop corresponding information-theoretic measure. Second, we investigate the suboptimal out-of-distribution (OOD) tokenization issue in large language models (LLM), and propose to improve OOD tokenization via a sparse optimal transport token translation.

Thesis Committee
J. Zico Kolter (Chair)
Graham Neubig
Ruslan Salakhutdinov
Lester Mackey (Microsoft Research, New England)

Additional Information

In Person and Zoom Participation.  See announcement.

Event Type: Thesis Proposals
Room Number: In Person and Virtual - ET
Building: Reddy Conference Room, Gates Hillman 4405 and Zoom
Speaker's Name: ZHILI FENG
Speaker Websitezhilif.github.io
Speaker's Professional Title: Ph.D. Student, Machine Learning Department, Carnegie Mellon University
Talk Title: Leveraging Information Theoretic Tools for Foundation Model Analysis
For More Informationstidle@andrew.cmu.edu
Affiliations: Machine Learning Department (MLD)
Organization(s): School of Computer Science