Teaching machines to understand sound is a new and exciting application for artificial intelligence. We are all familiar with speech and music recognition, but there is a far richer world of sound all around us. As the products we use as consumers have a greater sense of hearing they can start to unlock some exceptional human experiences. Each episode of the podcast we will be joined by expert guests with really interesting perspectives on the subject. We will be talking to people who will shed light on the importance and role of hearing and its applications, as well as people working at the cutting edge of acoustic science, data and machine learning.
For this episode of the podcast we had a fascinating conversation with Professor Catalin Grigoras and Cole Whitecotton from the National Center for Media Forensics within the University of Colorado. They shared a glimpse into their cutting-edge research into digital signals.
--
We were joined by two guests from the University of Colorado for the latest podcast. Professor Catalin Grigoras and Senior Professional Research Associate Cole Whitecotton. Both of our guests work for the National Center for Media Forensics within the university, with Professor Grigoras as the director.
We talked about the forensic trail from mains hum, metadata, deep fakes and the fascinating world of digital signal forensics. In a time of deepfakes, alternative facts and misinformation campaigns, we should recognize just how important media forensics is.
We first heard about their work on Electrical Network Frequency (ENF) which enables them to be able to confirm to the day, hour and minute of when a recording was made. But as this episode will highlight, their work goes far beyond ENF.
For more information on Catalin and Cole’s work, please visit the University of Colorado website: https://artsandmedia.ucdenver.edu/areas-of-study/national-center-for-media-forensics/about-the-national-center-for-media-forensics
In this episode of the Machine Listening Podcast I chat with Dr Chris Mitchell, founder and CEO of Audio Analytic. We talk about his journey to starting the company, the lessons he learned along the way and what it sounds like to canoe in the dark among a flock of moorhens.
Longer description:
My guest on the podcast this week is my colleague Dr Chris Mitchell, who founded Audio Analytic in 2010 following his PhD.
We talk about the work that led to the company being started, the process of raising awareness (and investment) while the world is talking about just speech recognition. He shares some amusing stories about how his experiences with sound have been shaped by the company and the future of machine listening that he finds exciting.
It was really enjoyable to sit down and chat with somebody who I know so well to learn more about his journey and experiences. I even learnt things that I didn’t previously know. We talked about the challenges of training a model to recognize soft drinks cans being opened and the unique experience of canoeing in the dark while disturbing a flock of sleeping moorhens.
For more information visit audioanalytic.com.
In this episode of the Machine Listening Podcast we chat with Professor Reddi from the Edge Computing Lab at Harvard University. We talked about a wide range of issues including the ‘AI tax’, sustainable AI, the drivers for running ML at the edge, the definition of tinyML, datasets and embedded ecosystems.
--
Professor Vijay Reddi on the AI tax and the drivers for tinyML
On the latest episode of the podcast I was joined by Professor Vijay Reddi. Professor Reddi is an Associate Professor in the John A. Paulson School of Engineering and Applied Sciences at Harvard University, where he focuses on mobile and edge-based computing systems and directs the university’s Edge Computing Lab. He also set up and ran the TinyML HarvardX course which teaches the fundamentals of machine learning and embedded devices.
In addition, he is a founding member of MLCommons, a non-profit focused on accelerating AI innovation, and a co-chair of MLPerf Inference, an organization that is responsible for fair and useful benchmarks for measuring the performance of ML systems.
During our chat we talked about his work looking at what he calls the AI tax – the overheads, the things that get in the way of building ML models that we have to incur. One significant challenge comes from data.
As Vijay says: “The elephant in the room is really how do you get the data, how do pre-process the data, how do you move that data into the neural network.”
In a wide ranging conversation we also talked about the drivers for running ML at the edge, how you define what ‘tiny’ means and what kind of embedded ecosystem will be required in the near future.
For more information on Vijay and his work, please visit the Harvard University website (https://scholar.harvard.edu/vijay-janapa-reddi/home). You can find the Harvard Edge Computing Lab here (https://edge.seas.harvard.edu/), MLCommons here (https://mlcommons.org/en/), and the TinyML course from HarvardX here (https://www.edx.org/professional-certificate/harvardx-tiny-machine-learning).
Audio Analytic’s Head of Research, Dr Çağdaş Bilen, talks to us about the challenges facing ML engineers when recognizing and enhancing sound. He also talked about his recent ICASSP presentation where he talked about the need to close the gap between probabilities and decisions in models.
Dr Çağdaş Bilen – How faithfully can ML models represent the real world?
My latest guest on the podcast is somebody I know very well - my colleague Dr Çağdaş Bilen, who is Head of Research at Audio Analytic.
Çağdaş has been with us since 2018 and he has contributed to and led our cutting edge research in a range of areas including a domain-specific loss function, post biasing, a powerful temporal decision engine and an evaluation metric called PSDS, which has quickly established itself as the standard evaluation framework for sound event detection tasks.
Recently Çağdaş was invited to speak as an industry expert at the prestigious IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) on the topic of probabilities versus decisions. In particular he used his speech to highlight the negative impact on users when machine learning researchers don’t think about end user impact and the huge gains that can be achieved through more sophisticated decision post-processing.
We talked about his presentation earlier this month at ICASSP, as well as many other areas of his work and the challenges facing ML researchers and engineers in this field.
You can watch Çağdaş’ keynote from ICASSP 2022 here.
You can read more about PSDS here, post-biasing here and loss function here.
For more information on DCASE click here and ICASSP click here.
One of the leading academics in the sound recognition community, Professor Mark Plumbley, talked to us about the challenges facing researchers, how the ML community needs to share negative results, and how he sees the future of sound recognition research shifting to something more user-centric.
On this episode of the Machine Listening podcast I was joined by two guests. Professor Mark Plumbley, Professor of Signal Processing at the UK’s University of Surrey, and Dr Cagdas Bilen, Head of Research at Audio Analytic, who joined me as a co-host.
In previous episodes we’ve talked about the acoustic spaces around us, the value of adding a sense of hearing to products as diverse as planetary robot explorers and the importance of realising that hearing isn’t just about the ears. This time we focus on the process of teaching machines to listen by talking to one of the pre-eminent academics in this field.
We talked about Mark’s interest in the field and his career before talking about the challenges that face researchers and how research needs to shift to starting with user needs.
When we touched on the DCASE Challenge, Mark said: “I’d like us to move away from the focus on ‘can we get the biggest numbers possible on some data measure’ but actually think about what the benefit is to people.”
He also called on the ML community to improve the visibility of the things that don’t work, so that researchers and students aren’t repeating the same mistakes.
“We are not working in the methodical way that say the medical research community might be where you need to register a study before you even try it. If it doesn’t work, everybody knows that it doesn’t work because you had to register it first. So you can’t sort of pretend that you didn’t try. Within the machine learning community as a whole we need to grapple with this and work out how do we improve the visibility of the things that don’t work, so it’s not just the tip of the iceberg that happen to work - maybe by chance - that you get to hear about.”
To learn more about Mark’s work you can visit the University of Surrey website and follow Mark on Twitter. To learn more about the research on using sound recognition to detect depression click here, and to learn more about DCASE here.
We spoke with the world-renowned, award-winning, and inspirational musician Dame Evelyn Glennie about how she uses her body as a resonating chamber to overcome the fact that she has been deaf since the age of 12.
“I started off thinking that you had to hear through the ears” – Dame Evelyn Glennie on being a better listener
My latest guest on the Machine Listening Podcast is double Grammy-winning, BAFTA-nominated Dame Evelyn Glennie. As a world-renowned solo percussionist, Evelyn has recorded over 40 albums, commissioned more than 200 pieces of music, performed the first percussion concerto in the history of the Proms at the Royal Albert Hall in 1992, and led 1,000 drummers during the opening ceremony of the London 2012 Olympics.
She is also Chancellor of Robert Gordon University in Aberdeen, Scotland, a popular public speaker and regularly composes music for film, TV and the theatre.
Since the age of 12 Evelyn has been deaf, which makes her achievements even more inspirational. Thanks to her early percussion teacher and an amazing career she has continued to hone her awareness of sound to such a degree that she describes her body as a resonating chamber.
This podcast is all about the subject of teaching machines to hear but with our latest guest it was a great opportunity to reflect on what it really means to listen.
We talked about her experiences, the way that she listens, the importance of being a better listener and the impressive ‘The Evelyn Glennie Collection’, which contains over 3,500 instruments and artefacts from her career so far.
If you aren’t familiar with her work, you can watch the London 2012 Olympic ceremony here:
https://youtu.be/4As0e4de-rI?t=1115
To learn more about Dame Evelyn Glennie visit evelyn.co.uk https://www.evelyn.co.uk/. Where you can listen and subscribe to Evelyn’s podcast https://www.evelyn.co.uk/theevelynglenniepodcast/, as well as learning more about ‘The Evelyn Glennie Collection’ https://www.evelyn.co.uk/the-evelyn-glennie-collection/.
We spoke with Professor Trevor Cox, who is a world-leading authority on acoustics, about a wide range of topics from his Guinness World Record, his journey to find the strangest acoustic locations on the planet, his Spinal Tap connection and the Clarity Challenge which looks to build a community around speech technology for people with hearing loss.
We spoke with Dr Baptiste Chide from Los Alamos National Laboratory to talk about his work on giving the NASA Perseverance Mars Rover the sense of hearing and the important role that sound will play in our exploration of the universe.
Plus, we got to talk about lasers, spacecraft and how there are two different speeds of sound on Mars, which would make music concerts a challenge.
Full episode description:
For our first episode of the Machine Listening Podcast we spoke with planetary scientist Dr Baptiste Chide from Los Alamos National Laboratory. Baptiste is a key member of the team who gave the NASA Perseverance Mars Rover the ability to listen to the alien planet.
Adding the sense of hearing to a spacecraft was previously considered a public outreach or PR opportunity by NASA as there weren’t any sounds on the red planet. However, on Perseverance they took their own sound source with them – the SuperCam laser, which zaps rocks.
Scientists including Dr Chide were able to prove that the sound of the laser sparks would give information on the rock types. As a result, two microphones were included to record the sounds of the laser as well as the space craft moving around the Martian surface, the atmosphere and even the first flight of the Ingenuity helicopter. All of which will open a new area of a space exploration. He hopes that NASA will use microphones on all the planets that have an atmosphere dense enough for sound to propagate, such as Venus and Saturn’s moon Titan.
Baptiste talked to us about the challenges of adding hardware that has to go through the harsh conditions of space flight as well as the realities of life on another planet. He talked about the interesting scientific discoveries that have been made by studying the sounds and what it felt like to be one of the first people to hear from the surface of Mars.
He said, “It was the first time that we were able to associate the sounds with the images we took on Mars. We had these landscapes, we had the headphones with the rumble of the wind and it was quite immersive. We were on Mars. It was very emotional.”
I hope that you enjoy this episode of The Machine Listening Podcast. It was fascinating for my colleague Arnoldas Jasonas and I to chat with Baptiste about his work and the future applications for the sense of hearing in ground-breaking space research.
If you want to learn more then please visit the Mars Mission page on the NASA website https://mars.nasa.gov/mars2020/. You can hear more sounds from Mars by visiting https://mars.nasa.gov/mars2020/multimedia/audio/. And if you want to know more about the Los Alamos National Laboratory, please visit https://www.lanl.gov/.
Welcome to The Machine Listening Podcast.
Teaching computers to understand sound is a new and exciting application for artificial intelligence. We are all familiar with speech and music recognition, but there is a far richer world of sound all around us.
As the products we use have a greater sense of hearing, they can start to unlock some exceptional human experiences.
That could involve enhancing what you can hear because of a challenging acoustic environment, helping to hear on your behalf, extending your sense of hearing to places where you aren’t present, anticipating what you might need based on the local context, or even helping you to remember what you heard.
But teaching machines to listen is challenging.
The human sense of hearing has its limitations and so rather than teaching machines to hear what we do, we have to better understand what and how they hear.
Each episode of the podcast, I will be joined by expert guests with really interesting perspectives on the subject. We’ll be talking to people who will shed light on the importance and role of sound and its applications, as well as people working at the cutting edge of acoustic science, data and machine learning.
I’m Dr Dominic Binks, I’m a musician, the VP of Technology at Audio Analytic and fascinated by sound. I’m really looking forward to speaking with our guests and I hope that you enjoy listening.
If you want to get in touch with the show please email podcast@machine-listening.com.