This talk examines speech recognition issue, comparing and contrasting them to what is known about human perception. With recent advances in Deep Learning, it is suggested that it is now achievable for Word Error Rates to be comparable to human listeners. This talk specifically highlights issues with accented, noisy speech, different speaking styles, multilingual speech recognition and more. And through demonstrations in comparison to human perception, there is still significant work in speech recognition research from the community.