Tuesday, August 6, 2024, 7pm
Events are crucial discourse elements in natural language, playing a vital role in semantic understanding due to their complex structures that interconnect various parts of discourse. Events interact with other discourse elements to form diverse structures. Extensive research has been conducted on analyzing them, primarily focusing on frame structures (examining semantic roles such as participants, time, and location) and various forms of anaphora involving multiple events and entities, such as event coreference, event schema (event sequence, script), and ellipsis. The rich interactive nature of events presents both challenges and opportunities. On one hand, predicting and analyzing event structures can be complex. For example, conducting a standard document level event slot filling, can involve multiple structure prediction tasks (e.g., event mention, coreference, arguments, and schemas). On the other hand, the interactions among these structures can be leveraged to enhance model predictions, or provide a lens to study the mechanisms of models. This thesis explores the complexities and benefits of such interactions across different data availability scenarios, developing prediction methods including direct supervised training, crowdsourced event datasets, and studying automatically formed mechanisms related to them.
In the first part, we present empirical results analyzing event semantics with expert-annotated task-specific annotated datasets. We start by introducing methods studying isolated structures, such as event mention prediction, pair-wise event coreference, and event sequencing. We then present approaches to solve problems involving multiple structures, using multi-step or joint learning methods, such as joint coreference and sequencing, slot filling and verb phrase ellipsis. Recognizing the high cost of scaling expert-annotated datasets, the second part of this thesis explores methods to increase data availability through crowdsourcing and indirect supervision signals. An intriguing outcome of these approaches is their ability to reveal previously unspecified interactions between different textual structures. A key contribution in this area is our LLM360 language model project, which shares intermediate checkpoints throughout a model’s training process. We demonstrate the project’s utility for interpretability analysis, using the complex anaphora task of Winograd schemas as a case study.
This thesis demonstrates methods to address and utilize the complex nature of events. We find that scaling up data size leads to more iterations of such structures “automatically” appearing during analysis. Looking ahead, this work opens avenues for developing more sophisticated models that better capture relationships between event structures and for exploring large language models’ potential in understanding complex event semantics. Additionally, our interpretability analysis, particularly through LLM360, paves the way for investigating how these models processevent semantics internally. This research could lead to more transparent and explainable AI systems, advancing our understanding of complex language processing.
Thesis Committee:
Teruko Mitamura (Chair)
Eduard Hovy
Taylor Berg-Kirkpatrick (University of California San Diego)
Vicent Ng (The University of Texas at Dallas)
Additional Information
Zoom Participation. See announcement.
Event Type: Thesis Orals
Room Number: Virtual Presentation - ET
Building: Remote Access - Zoom
Speaker's Name: ZHENGZHONG (HECTOR) LIU
Speaker Website: hunterhector.github.io
Speaker's Professional Title: Ph.D. Candidate, Language Technologies Institute, Carnegie Mellon University
Talk Title: Diving Deep into Event Semantics
Event Poster Title: Poster
Event Poster URL: www.cs.cmu.edu…
For More Information: StaceyYoung@cmu.edu
Affiliations: Language Technologies Institute (LTI)
Organization(s): School of Computer Science