Thursday, April 20, 2023, 12 – 1pm
Algorithmic reasoning requires capabilities which are most naturally understood through recurrent models of computation, like the Turing machine. However, Transformer models, while lacking recurrence, are able to perform such reasoning using far fewer layers than the number of reasoning steps. This raises the question: what solutions are these shallow and non-recurrent models finding? In this talk, we will formalize reasoning in the setting of automata, and show that the computation of an automaton on an input sequence of length T can be replicated exactly by Transformers with o(T) layers, which we call "shortcuts". We provide two constructions, with O(log T) layers for all automata and O(1) layers for solvable automata. Empirically, our results from synthetic experiments show that shallow solutions can also be found in practice.
—
Bingbin Liu is a fourth-year PhD student at the Machine Learning Department of Carnegie Mellon University, co-advised by Pradeep Ravikumar and Andrej Risteski. Her research focuses on the theoretical understanding of self-supervised learning and unsupervised learning, often motivated by findings in vision and language.
The AI Seminar is generously sponsored by SambaNova Systems.
In Person and Zoom Participation. See announcement.
Event Type: Seminars
Room Number: In Person and Virtual - ET
Building: ASA Conference Room, Gates Hillman 6115 and Zoom
Speaker's Name: BINGBIN LIU
Speaker Website: clarabing.github.io
Speaker's Professional Title: Ph.D. Student, Machine Learning Department, Carnegie Mellon University
Talk Title: Thinking Fast with Transformers: Algorithmic Reasoning via Shortcuts
For More Information: ashert@cs.cmu.edu
Affiliations: Computer Science Department (CSD), Language Technologies Institute (LTI), Machine Learning Department (MLD)
Event Website Title: Series Website
Event Website URL: www.cs.cmu.edu…