Tuesday, December 6, 2022, 8:35am – 1:10pm
► 8:35 - 9:15 am → Posner Hall 153 and Zoom
Simbot Navigation and Interaction
— Adhokshaja Madhwaraj, Jessica Zhong, Kushagra Mahajan, Malaika Vijay,
Sai Vishwas Padigi, Vineeth Reddy Vatti
Embodied task-completion agents are intelligent agents that can perceive, navigate, and manipulate objects in an environment. We developed a bot for the SimBot challenge in conversational embodied task completion, where users can converse with and guide the bot to complete tasks in a 3D environment. The Navigation and Interaction thrust of the SimBot project focuses on scene understanding and action planning to perform actions that satisfy a user’s instruction. Our primary contributions include an Image Segmentation model to identify objects of interest and an Action Planning module capable of generating a logical sequence of actions to execute on the simulator.
► 9:15 - 9:55 am → Posner Hall 153 and Zoom
Simbot Dataset Collection: An Embodied Vision-Audio-Navigation Task
—Bharani Ujjaini Kempaiah, Linyi Li
Intelligent robotic agents that operate in human spaces must be capable of realizing and executing instructions that are conveyed in natural language via speech. Most existing vision and language benchmarks contain instructions in natural language via text, and subsequently, models trained on these datasets also rely heavily on text as an intermediate medium to achieve grounding. In this work, we aim to enhance the existing ALFRED benchmark by crowd-sourcing speech annotations for existing ALFRED demonstrations of household tasks. We build web interfaces to qualify crowd workers before collecting speech annotations online. With these speech annotations, we believe the research community can develop models that can power robots to achieve audio-visual grounding similar to how humans interact in the real world.
► 11:50 am - 1:10 pm → Posner Hall 153 and Zoom
Learn-to-Race: A Multimodal Control Environment for Autonomous Racing
— Kevin Chian, Arav Agarwal, Sidharth Kathpal, Yujun Qin, Tanay Gangey
Autonomous racing is a sub-field of autonomous driving which has been studied considerably less than urban driving. To further research in this area, we implement a set of extensions spanning reinforcement learning (RL), Computer Vision (CV), and robotics interfaces upon the Learn-to-Race framework, which itself runs on top of a racing simulator. Our project also has a software development component that involves creating interfaces to connect to an actual vehicle using ROS. These research and development thrusts are crucial to designing safe and fast autonomous agents, as failures in real-life are exceptionally costly. The safe policies learned by the autonomous racers can be generalized to the real world to have safer and faster autonomous driving agents.
The MCDS Capstone projects continue December 6, 7, 8, 12 and 13.
Please join in as your schedule permits.
Event Type: Project Presentations
Room Number: In Person and Virtual - ET
Event Poster Title: Full Schedule of Daily Presentations
Event Poster URL: mcds-cmu.github.io…
For More Information: ahan2@andrew.cmu.edu
Affiliations: Language Technologies Institute (LTI)
Organization(s): MCDS