Welcome to the VizWiz podcast, where we explore two questions: "what are visual challenges blind people encounter in their daily lives?" and "what are technological advancements helping blind people to live more independently?" The hosts are a diverse team of researchers from academia and industry with a shared passion around developing more inclusive artificial intelligence (AI) solutions. We interview the world’s leading AI researchers, technology advocates from the blind community, and industry specialists to raise awareness about current access technologies and help inspire their future.
Brief summary of the episode Saqib Shaikh, Will Butler, and Karthik Kannan share about their experiences developing visual interpretation products and services for blind and low vision users.
Questions asked in the episode
Guest bios
Saqib Shaikh is an Engineer Manager at Microsoft, where he founded Seeing AI - an app which enables someone who is visually impaired to hold up their phone, and hear more about the text, people, and objects in their surroundings.
Will Butler is the Chief Experience Officer at Be My Eyes, a free app with one of the largest online communities, that supports visually impaired individuals to get free, on-demand, live video support from around 4 million volunteers and companies. Will also has hosted two podcasts on the topic of vision loss and accessibility.
Karthik Kannan is the co-founder and chief technology officer of Envision, a company that provides technology to help people with visual impairments in their daily lives. His company builds an app and smart glasses which helps people with visual impairments learn about their surroundings including to read text and recognize faces.
Samreen Anjum is a PhD student at University of Colorado Boulder. Her research focuses on computer vision and its applications in the fields of biomedical sciences and assistive technologies as well as designing systems that enable collaborations between humans and machines.
Links to resources mentioned * https://vizwiz.org/workshops/2022-workshop/
Brief summary of the episode
Stephanie Enyart, Robin Christopherson, and Daniel Kish share about their experiences advocating for better visual interpretation technology for blind and low vision users as well as their experiences—as people with visual impairments—using such technologies.
Questions asked in the episode
Guest bios
Stephanie Enyart is the Chief Public Policy and Research Officer at the American Foundation for the Blind. Stephanie serves as a strategic leader in developing policy that benefits people who are blind in education, employment, aging, and the intersectional issues of technology and transportation.
Robin Christopherson is a co-founder and Head of Digital Inclusion at AbilityNet. His work has led to accessibility improvements in many organizations spanning industry, government, and universities. Robin also has served as an expert technical witness around assistive technology in software, systems and websites.
Daniel Kish is the President of World Access for the Blind. He is a world leader in perceptual navigation and ecolocation, through which he has developed his own method of generating vocal clicks and using echoes to identify his surroundings and navigate.
Abigale Stangl is a CRA/NSF Computing Innovation Fellow at the University of Washington. Her research lies at the intersection of human-computer interaction, non-visual accessibility, and data privacy and ownership.
Links to resources mentioned
Brief summary of the episode Stephanie Enyart, James Coughlan, and Karthik Kannan share very diverse perspectives around the development of visual interpretation technologies to meet the interests and needs of people with vision impairments.
Questions asked in the episode
Guest bios
Stephanie Enyart is the Chief Public Policy and Research Officer at the American Foundation for the Blind. Stephanie serves as a strategic leader in developing policy that benefits people who are blind in education, employment, aging, and the intersectional issues of technology and transportation.
James Coughlan is a Senior Scientist at the Smith-Kettlewell Eye Research Institute, with a PhD in Physics from Harvard University. James has been at Smith-Kettlewell since 1998 and over this time has developed a wide array of impactful technologies for the blind and low-vision community.
Karthik Kannan is the co-founder and chief technology officer of Envision, a company that provides technology to help people with visual impairments in their daily lives. His company builds an app and smart glasses which helps people with visual impairments learn about their surroundings including to read text and recognize faces.
Danna Gurari is an Assistant Professor at University of Colorado Boulder where she also leads the Image and Video Computing research group.
Links to resources mentioned * https://vizwiz.org/workshops/2022-workshop/
Brief summary of the episode Daniel Kish, Andrew Howard, and Will Butler share very diverse perspectives around the development of visual interpretation technologies to meet the interests and needs of people with vision impairments.
Questions asked in the episode
Guest bios
Daniel Kish is the President of World Access for the Blind. He is a world leader in perceptual navigation and ecolocation, through which he has developed his own method of generating vocal clicks and using echoes to identify his surroundings and navigate.
Andrew Howard is a Senior Staff Software Engineer at Google Research, with a PhD in Computer Science from Columbia University. Andrew is most well-known for his work in mobile-friendly deep learning models. Starting with MobileNets, then MobileNetsV2, then MobileNetsV3, and also MnasNets, his work has been broadly adopted in deep learning packages like PyTorch and Tensorflow as well as across a host of mobile phone platforms and apps.
Will Butler is the Chief Experience Officer at Be My Eyes, a free app with one of the largest online communities, that supports visually impaired individuals to get free, on-demand, live video support from around 4 million volunteers and companies. Will also has hosted two podcasts on the topic of vision loss and accessibility.
Ed Cutrell is a Senior Principal Research Manager at Microsoft Research (MSR), where he leads the MSR Ability Team, a group of researchers focused on innovating new technologies for people with a range of disabilities.
Links to resources mentioned * https://vizwiz.org/workshops/2022-workshop/
Brief summary of the episode Marcus Rohrbach, Andrew Howard, and James Coughlan share about their experiences developing state-of-art research in visual interpretation algorithms and systems.
Questions asked in the episode
Guest bios Marcus Rohrbach is a Research Scientist at Meta AI Research, with a PhD from the Max Planck Institute for Informatics. Marcus is most well-known for his work at the intersection of computer vision and natural language processing. Over his career, he has driven key progress in visual question answering, language grounding, and generating descriptions about image and videos, in particular movies - all of which is highly relevant for the blind/low-vision community.
Andrew Howard is a Senior Staff Software Engineer at Google Research, with a PhD in Computer Science from Columbia University. Andrew is most well-known for his work in mobile-friendly deep learning models. Starting with MobileNets, then MobileNetsV2, then MobileNetsV3, and also MnasNets, his work has been broadly adopted in deep learning packages like PyTorch and Tensorflow as well as across a host of mobile phone platforms and apps.
James Coughlan is a Senior Scientist at the Smith-Kettlewell Eye Research Institute, with a PhD in Physics from Harvard University. James has been at Smith-Kettlewell since 1998 and over this time has developed a wide array of impactful technologies for the blind and low-vision community.
Daniela Massiceti is a machine learning researcher at Microsoft Research. Her research focuses on the intersection of ML and human-computer interaction. She is primarily interested in ML systems that learn and evolve with human input, so called “teachable” systems, giving users the power to completely customise their AI experiences – from personalised assistive tools for people who are blind/low-vision, to personalised avatars in the metaverse.
Links to resources mentioned * https://vizwiz.org/workshops/2022-workshop/
Brief summary of the episode Robin Christopherson, Saqib Shaikh, and Marcus Rohrbach share very diverse perspectives around the development of visual interpretation technologies to meet the interests and needs of people with vision impairments.
Questions asked in the episode
Guest bios
Robin Christopherson is a co-founder and Head of Digital Inclusion at AbilityNet. His work has led to accessibility improvements in many organizations spanning industry, government, and universities. Robin also has served as an expert technical witness around assistive technology in software, systems and websites.
Saqib Shaikh is an Engineer Manager at Microsoft, where he founded Seeing AI - an app which enables someone who is visually impaired to hold up their phone, and hear more about the text, people, and objects in their surroundings.
Marcus Rohrbach is a Research Scientist at Meta AI Research, with a PhD from the Max Planck Institute for Informatics. Marcus is most well-known for his work at the intersection of computer vision and natural language processing. Over his career, he has driven key progress in visual question answering, language grounding, and generating descriptions about image and videos, in particular movies - all of which is highly relevant for the blind/low-vision community.
Danna Gurari is an Assistant Professor at University of Colorado Boulder where she also leads the Image and Video Computing research group.
Links to resources mentioned * https://vizwiz.org/workshops/2022-workshop/