A weekly show answering the question, "How is the Internet holding up this week?" Watch each week to understand the latest outage trends across global ISPs, public cloud providers, collaboration app networks, and edge networks like CDNs, DNS, SECaaS, etc.
Backend-related outages were somewhat of a theme during the first half of 2025 and this trend seems to be continuing.
In recent weeks, we’ve seen several service disruptions that appeared connected to backend issues, affecting organizations including a social media platform, an Australian bank, and a major airline.
These incidents offer valuable insights into the typical characteristics of backend outages and strategies for spotting them.
Tune in now to hear more about these events and the distinct anatomy of a backend outage.
———
CHAPTERS
00:00 Intro
01:01 About Backend Outages
04:18 Anthropic’s Claude Outage
08:44 Service Disruption
12:54 Alaska Airlines Outage
14:17 Commonwealth Bank Outage
15:42 Gong Service Disruption
17:25 Outage Trends: By the Numbers
19:43 Get in Touch
———
Explore Anthropic’s Claude outage and the Alaska Airlines outage further in the ThousandEyes platform (no login required):
Claude: https://acshghvfamioplwqigifrycvlsqesoxb.share.thousandeyes.com
Alaska Airlines: https://annocixygjjriicwjupcporbjqblgkym.share.thousandeyes.com/
For additional insights, check out the links below:
The Internet Outage Survival Kit: https://www.thousandeyes.com/resources/the-internet-outage-survival-kit?utm_source=wistia&utm_medium=referral&utm_campaign=fy26q1_internetreport_q1fy26ep6_podcast
The Guide to Next-generation Assurance: https://www.thousandeyes.com/resources/guide-to-next-generation-assurance-ebook?utm_source=wistia&utm_medium=referral&utm_campaign=fy26q1_internetreport_q1fy26ep6_podcast
———
Want to get in touch?
If you have questions, feedback, or guests you would like to see featured on the show, send us a note at InternetReport@thousandeyes.com. Or follow us on LinkedIn or X: @thousandeyes
———
ABOUT THE INTERNET REPORT
This is The Internet Report, a podcast uncovering what’s working and what’s breaking on the Internet—and why.
Tune in to hear ThousandEyes’ Internet experts dig into some of the most interesting outage events from the past couple weeks, discussing what went awry—was it the Internet, or an application issue?
Plus, learn about the latest trends in ISP outages, cloud network outages, collaboration network outages, and more.
Catch all the episodes on YouTube or your favorite podcast platform:
Apple Podcasts: https://podcasts.apple.com/us/podcast/the-internet-report/id1506984526
Spotify: https://open.spotify.com/show/5ADFvqAtgsbYwk4JiZFqHQ?si=00e9c4b53aff4d08&nd=1&dlsi=eab65c9ea39d4773
SoundCloud: https://soundcloud.com/ciscopodcastnetwork/sets/the-internet-report
Just as a detective must gather clues and consider all the evidence, ITOps teams investigating an outage must consider all available data points in context to understand where system failures likely occurred. Recent outages that impacted Starlink and Google Maps SDKs powerfully illustrated this point.
Tune in to learn more about these incidents and also get an update on the September Red Sea cable cuts.
———
CHAPTERS
00:00 Intro
00:53 Starlink Outage
03:39 Google Maps Outage
06:50 Update: Red Sea Cable Cuts
13:26 Outage Trends: By the Numbers
14:34 Get in Touch
———
Want to get in touch?
If you have questions, feedback, or guests you would like to see featured on the show, send us a note at InternetReport@thousandeyes.com. Or follow us on LinkedIn or X: @thousandeyes
———
ABOUT THE INTERNET REPORT
This is The Internet Report, a podcast uncovering what’s working and what’s breaking on the Internet—and why.
Tune in to hear ThousandEyes’ Internet experts dig into some of the most interesting outage events from the past couple weeks, discussing what went awry—was it the Internet, or an application issue?
Plus, learn about the latest trends in ISP outages, cloud network outages, collaboration network outages, and more.
Catch all the episodes on YouTube or your favorite podcast platform:
Apple Podcasts: https://podcasts.apple.com/us/podcast/the-internet-report/id1506984526
Spotify: https://open.spotify.com/show/5ADFvqAtgsbYwk4JiZFqHQ?si=00e9c4b53aff4d08&nd=1&dlsi=eab65c9ea39d4773
SoundCloud: https://soundcloud.com/ciscopodcastnetwork/sets/the-internet-report
Gain insights on the Red Sea subsea cable cuts; connectivity issues in China; and other recent service disruptions at Verizon, Google, and Mailchimp.
———
CHAPTERS
00:00 Intro
01:00 Red Sea Cable Cuts
05:35 Connectivity Issues in China
09:20 Verizon
10:46 Google
12:08 Mailchimp
13:05 Outage Trends: By the Numbers
15:07 Get in Touch
———
Want to get in touch?
If you have questions, feedback, or guests you would like to see featured on the show, send us a note at InternetReport@thousandeyes.com. Or follow us on LinkedIn or X: @thousandeyes
———
ABOUT THE INTERNET REPORT
This is The Internet Report, a podcast uncovering what’s working and what’s breaking on the Internet—and why.
Tune in to hear ThousandEyes’ Internet experts dig into some of the most interesting outage events from the past couple weeks, discussing what went awry—was it the Internet, or an application issue?
Plus, learn about the latest trends in ISP outages, cloud network outages, collaboration network outages, and more.
Catch all the episodes on YouTube or your favorite podcast platform:
Apple Podcasts: https://podcasts.apple.com/us/podcast/the-internet-report/id1506984526
Spotify: https://open.spotify.com/show/5ADFvqAtgsbYwk4JiZFqHQ?si=00e9c4b53aff4d08&nd=1&dlsi=eab65c9ea39d4773
SoundCloud: https://soundcloud.com/ciscopodcastnetwork/sets/the-internet-report
Dive into the world of Border Gateway Protocol (BGP)—the backbone of the Internet—and explore everything from BGP zombies to BGP monitoring best practices.
Tune in for this special conversation with Lefteris Manassakis and The Internet Report team. A seasoned researcher and network engineer, Lefteris Manassakis co-founded Code BGP, which is now part of Cisco ThousandEyes. He currently serves as a Software Engineering Technical Leader at ThousandEyes, focusing on BGP monitoring. To learn more, follow Lefteris on LinkedIn (https://www.linkedin.com/in/manassakis/) or visit his website (https://manassakis.net/)
———
CHAPTERS
00:00 Intro
01:11 What Is BGP?
07:05 May 2023 Incident
17:16 Challenges of BGP Monitoring
19:16 BGP Zombies
33:45 Get in Touch
———
For additional BGP insights, check out these links:
Monitoring Root DNS Prefixes: https://www.thousandeyes.com/blog/monitoring-root-dns-prefixes?utm_source=youtube&utm_medium=referral&utm_campaign=fy26q1_internetreport_q1fy26ep3_podcast
BGP Zombies Show Up Regularly: https://www.thousandeyes.com/blog/bgp-zombies-show-up-regularly?utm_source=youtube&utm_medium=referral&utm_campaign=fy26q1_internetreport_q1fy26ep3_podcast
Monitor BGP Routes to and From Your Network: https://www.thousandeyes.com/solutions/bgp-and-route-monitoring?utm_source=youtube&utm_medium=referral&utm_campaign=fy26q1_internetreport_q1fy26ep3_podcast
———
Want to get in touch?
If you have questions, feedback, or guests you would like to see featured on the show, send us a note at InternetReport@thousandeyes.com. Or follow us on LinkedIn or X: @thousandeyes
———
ABOUT THE INTERNET REPORT
This is The Internet Report, a podcast uncovering what’s working and what’s breaking on the Internet—and why.
Tune in to hear ThousandEyes’ Internet experts dig into some of the most interesting outage events from the past couple weeks, discussing what went awry—was it the Internet, or an application issue?
Plus, learn about the latest trends in ISP outages, cloud network outages, collaboration network outages, and more.
Catch all the episodes on YouTube or your favorite podcast platform:
Apple Podcasts: https://podcasts.apple.com/us/podcast/the-internet-report/id1506984526
Spotify: https://open.spotify.com/show/5ADFvqAtgsbYwk4JiZFqHQ?si=00e9c4b53aff4d08&nd=1&dlsi=eab65c9ea39d4773
SoundCloud: https://soundcloud.com/ciscopodcastnetwork/sets/the-internet-report
When we looked at outages from the first half of 2025, we saw distinct patterns in how distributed systems failed. Consistent with our ongoing outage analysis, our data revealed a rise in subtle functional failures and service degradations where symptoms often seem disconnected from their root causes.
Tune in to hear more about our findings from 2025 outages so far, or use the chapters below to jump to the sections that most interest you.
CHAPTERS
00:00 Intro
00:46 The Impact of Agile Development
04:40 Hidden Functional Failures
08:04 Outages With Cascading Effects
11:52 Configuration-related Outages
12:39 Backend Issues
16:55 By the Numbers: 2025 Outage Trends
18:11 Get in Touch
———
For additional insights, check out the links below:
The Internet Report’s latest blog post: https://www.thousandeyes.com/blog/internet-report-outage-patterns-in-2025?utm_source=wistia&utm_medium=referral&utm_campaign=fy26q1_internetreport_q1fy26ep2_podcast
The Guide to Next-generation Assurance: https://www.thousandeyes.com/resources/guide-to-next-generation-assurance-ebook?utm_source=wistia&utm_medium=referral&utm_campaign=fy26q1_internetreport_q1fy26ep2_podcast
Internet Outage Map: https://www.thousandeyes.com/outages?utm_source=wistia&utm_medium=referral&utm_campaign=fy26q1_internetreport_q1fy26ep2_podcast
Internet Outages Timeline: https://www.thousandeyes.com/resources/internet-outages-timeline?utm_source=wistia&utm_medium=referral&utm_campaign=fy26q1_internetreport_q1fy26ep2_podcast
Internet Outage Survival Kit: https://www.thousandeyes.com/resources/the-internet-outage-survival-kit?utm_source=wistia&utm_medium=referral&utm_campaign=fy26q1_internetreport_q1fy26ep2_podcast
———
Want to get in touch?
If you have questions, feedback, or guests you would like to see featured on the show, send us a note at InternetReport@thousandeyes.com. Or follow us on LinkedIn and X at @thousandeyes.
———
ABOUT THE INTERNET REPORT
This is The Internet Report, a podcast uncovering what’s working and what’s breaking on the Internet—and why.
Tune in to hear ThousandEyes’ Internet experts dig into some of the most interesting outage events from the past couple weeks, discussing what went awry—was it the Internet, or an application issue?
Plus, learn about the latest trends in ISP outages, cloud network outages, collaboration network outages, and more.
Catch all the episodes on YouTube or your favorite podcast platform:
Apple Podcasts: https://podcasts.apple.com/us/podcast/the-internet-report/id1506984526
Spotify: https://open.spotify.com/show/5ADFvqAtgsbYwk4JiZFqHQ?si=00e9c4b53aff4d08&nd=1&dlsi=eab65c9ea39d4773
SoundCloud: https://soundcloud.com/ciscopodcastnetwork/sets/the-internet-report
From robotaxis to surgical robots, the innovations we’ve witnessed in industrial IoT (IIoT) and robotics technology wouldn't have been possible without advancements in networking.
Tune into this episode to learn more and explore why smooth digital experiences are crucial in the IIoT and robotics space.
CHAPTERS
00:00 Intro
00:55 Networking Advancements & IIoT Innovation
02:38 Assuring IIoT Networks
03:59 Impact of Disruptions
07:30 Monitoring Challenges
10:20 What’s Next
15:08 Get in Touch
———
For additional insights, check out this Guide to Next-generation Assurance: https://www.thousandeyes.com/resources/guide-to-next-generation-assurance-ebook?utm_source=wistia&utm_medium=referral&utm_campaign=fy26q1_internetreport_q1fy26ep1_podcast
———
Want to get in touch?
If you have questions, feedback, or guests you would like to see featured on the show, send us a note at InternetReport@thousandeyes.com. Or follow us on LinkedIn or X: @thousandeyes
———
ABOUT THE INTERNET REPORT
This is The Internet Report, a podcast uncovering what’s working and what’s breaking on the Internet—and why.
Tune in to hear ThousandEyes’ Internet experts dig into some of the most interesting outage events from the past couple weeks, discussing what went awry—was it the Internet, or an application issue?
Plus, learn about the latest trends in ISP outages, cloud network outages, collaboration network outages, and more.
Catch all the episodes on YouTube or your favorite podcast platform:
Apple Podcasts: https://podcasts.apple.com/us/podcast/the-internet-report/id1506984526
Spotify: https://open.spotify.com/show/5ADFvqAtgsbYwk4JiZFqHQ?si=00e9c4b53aff4d08&nd=1&dlsi=eab65c9ea39d4773
SoundCloud: https://soundcloud.com/ciscopodcastnetwork/sets/the-internet-report
As AI transforms IT infrastructure, it’s also reshaping what it takes for IT operations teams to assure performance and maintain quality digital experiences.
In this episode, we’ll explore the new challenges facing ITOps teams as AI becomes more integrated into IT environments, covering key digital resilience strategies and important considerations.
CHAPTERS
00:00 Intro
00:54 AI & IT Infrastructure
03:15 Assuring Performance on Your AI Journey
04:55 Distributed Architecture
07:49 Digital Resilience
10:21 Catching Issues in the AI Era
11:47 Performance Problems
13:57 AI Readiness: A Journey, Not a Destination
15:07 Get in Touch
———
For additional insights, check out this Guide to Next-generation Assurance: https://www.thousandeyes.com/resources/guide-to-next-generation-assurance-ebook?utm_source=transistor&utm_medium=referral&utm_campaign=fy25q4_internetreport_q4fy25ep5_podcast
———
Want to get in touch?
If you have questions, feedback, or guests you would like to see featured on the show, send us a note at InternetReport@thousandeyes.com. Or follow us on LinkedIn or X @thousandeyes.
This is The Internet Report, where we uncover what’s working and what’s breaking on the Internet—and why. In this episode, we unpack four notable outages that impacted WhatsApp, Zscaler, Salesforce, and Facebook, which all appear to have a common theme. Join our co-hosts Mike Hicks, Principal Solutions Analyst at ThousandEyes, and Chris Villemez, Technical Marketing Engineer at ThousandEyes, as they walk through each incident to understand what happened and discuss how network professionals can attempt to mitigate these types of scenarios in the future.
FURTHER READING Facebook Outage Analysis → https://www.thousandeyes.com/blog/facebook-outage-analysis
We're back! 00:00 Welcome: This is The Internet Report, where we uncover what’s working and what’s breaking on the Internet—and why. On this episode, our newest host, Chris Villemez, is joined by Kemal Sanjta to discuss a BGP-related incident that took down Twitter for many users around the globe on March 28th. 00:36 Under the Hood: Chris Villemez and Kemal Sanjta leverage their extensive operations experience managing the networks of large-scale SaaS, IoT, and cloud providers to analyze this incident using the ThousandEyes platform. They examine the scope of the outage, review the specific BGP changes that resulted in the outage, and discuss what enterprises can do when they’re experiencing a similar BGP hijack or route leak. Sharelinks: Single agent (Manchester) test: https://anislusvvn.share.thousandeyes.com/ Multi-agent global test showing BGP changes: https://axntbxntyk.share.thousandeyes.com/ 31:00 Outro: We've been on a bit of a break for the past several months as things were relatively quiet on the Internet front and for the foreseeable future we'll be a bit reactive in our episodes, when something major happens trust we'll be here. Questions? Feedback? Have an idea for a guest? Send us an email at internetreport@thousandeyes.com
This is The Internet Report, where we uncover what’s working and what’s breaking on the Internet—and why. On today’s episode, our newest host and Technical Marketing Engineer, Chris Villemez, is joined by Kemal Sanjta, Principal Engineer, to dive into the details of the recent AWS outages from December 7th, 10th and 15th. They’ll walk through what ThousandEyes saw from its fleet of vantage points, as well as share some insight into what enterprises can learn from these incidents to build resilient cloud architectures.
00:00 Welcome: This is The Internet Report, where we uncover what’s working and what’s breaking on the Internet—and why. 00:15 Headlines: Today we’re going to do a thorough analysis of the major Facebook outage that took place yesterday, Monday, October 4. I’m joined by Gustavo Ramos, ThousandEyes’ in-house expert on Network Engineering. ThousandEyes Blog: https://www.thousandeyes.com/blog/facebook-outage-analysis Analysis from Facebook: https://engineering.fb.com/2021/10/05/networking-traffic/outage-details/ 1:17 Under the Hood: We'll walk through the sequence of events that led to this outage, understand what went wrong (and what actions may have made the situation worse), and what lessons we can all learn from this outage. 25:40 Outro: We've been on a bit of a break for the past several months as things were relatively quiet on the Internet front and for the foreseeable future we'll be a bit reactive in our episodes, when something major happens trust we'll be here. Questions? Feedback? Have an idea for a guest? Send us an email at internetreport@thousandeyes.com
00:00 Welcome: This is The Internet Report, where we uncover what’s working and what’s breaking on the Internet—and why. 00:08 Headlines: Today, Mike Hicks (Principal Solutions Analyst, ThousandEyes) and I discuss a recent BGP routing incident that had intermittent impacts on Amazon’s services, including Amazon.com and AWS compute resources, during a five-hour period on July 12. 01:04 Under the Hood: When we look into BGP routing at the time, we can see multiple BGP path changes due to a service provider erroneously inserting themselves into the path for a large number of Amazon routes. Watch this episode to see how the BGP incident led to significant packet loss, resulting in service disruption for some Amazon and AWS users. We also discuss why enterprises need to have continuous oversight of the paths their traffic takes over the Internet. 17:58 Outro: Questions? Feedback? Have an idea for a guest? Send us an email at internetreport@thousandeyes.com
This is The Internet Report, where we uncover what’s working and what’s breaking on the Internet—and why. I’m joined today by Mike Hicks, principal solutions analyst here at ThousandEyes, to cover the outage of Akamai’s DNS service. The outage, which occurred on July 22nd around 3:38 PM UTC (8:38AM PT), struck during the course of business hours in Europe and North America, resulting in widespread impacts to applications and services hosted within Akamai servers. The outage itself was short-lived and was resolved roughly one hour after the outage began.
In this episode, we examine the customer impact, the relationship between DNS and CDNs, and what enterprises should take away from the incident. We also discuss the question of when it might make sense to invest in DNS or CDN redundancy—and when it is, frankly, overkill. Watch this week’s episode to hear our take, and as always let us know on Twitter what you think.
00:00 Welcome:This is The Internet Report, where we uncover what’s working and what’s breaking on the Internet—and why. 00:13 Headlines: Today, Kemal and I unpack an interesting BGP incident, in which a large-scale route leak briefly altered traffic patterns across the Internet. 00:58 Under the Hood: The incident began on Thursday, June 3rd at around 10:24 UTC, and resulted in a significant spike in packet loss that was noticeable in ThousandEyes tests. While this packet loss resolved within the hour (at around 10:48 UTC), we observed some interesting routing changes during this window—as traffic was diverted to a Russian telecom provider that had not previously been in the path. Watch this episode as we explore how this network provider managed to get itself into the routing paths of many major services, and why network visibility is so important to recognize these types of incidents in which your site may still be reachable but your traffic is being sent through an unexpected network. 20:45 Outro: Questions? Feedback? Have an idea for a guest? Send us an email at internetreport@thousandeyes.com
This is The Internet Report, where we uncover what’s working and what’s breaking on the Internet—and why. I’m joined by ThousandEyes’ BGP expert, Kemal Sanjta, to review the June 16th outage of Prolexic Routed, a DDoS Mitigation Service operated by Akamai. According to a statement from Akamai, the outage was not due to a DDoS attack or system update, but instead a routing table limitation that was inadvertently exceeded.
In this episode, Kemal and I analyzed what happened and how customers of Akamai Prolexic who had automated failover mechanisms in place were able to recover more quickly than those that had to manually switch over to other providers. Watch this episode to learn more about this outage, and how different operational processes resulted in very different service outcomes.
00:00 Welcome: This is The Internet Report, where we uncover what’s working and what’s breaking on the Internet—and why. 00:12 Headlines: Today, I’m joined by Hans Ashlock, Director of Technology & Innovation at ThousandEyes, to unpack today’s major outage at Fastly, a popular CDN provider. 3:46 Under the Hood: Today, I’m joined by Hans Ashlock, Director of Technology & Innovation at ThousandEyes, to unpack today’s major outage at Fastly, a popular CDN provider. The widespread outage occurred around 9:50 UTC, about 5:50 am ET, and mostly impacted users across Europe and Asia due to the timing. he outage lasted approximately one hour until 10:50 UTC, yet residual impacts were felt beyond that. Today’s outage is a good example of the importance of having outside-in visibility not just across your app, but also to your app’s edge and all its dependent services. 39:05 Outro: Questions? Feedback? Have an idea for a guest? Send us an email at internetreport@thousandeyes.com
This is The Internet Report, where we uncover what’s working and what’s breaking on the Internet—and why. I’m joined today by Mike Hicks, Principle Solution Analyst at ThousandEyes, to cover two recent application-related outages. The first occurred on May 19th around 12:50 UTC at Coinbase—a well-known cryptocurrency exchange. Around the time that news broke saying that the Chinese government would be imposing strict regulation on cryptocurrencies, users attempting to execute transactions were unable to access the application. From the ThousandEyes platform we were able to see a drop in availability around this time as well as increased load times (which in some cases resulted in timeout errors).
The second outage happened on May 20th around 17:35 UTC at Slack—an enterprise collaboration platform. While the outage was resolved within 90 minutes, it occurred during normal US business hours, making it particularly disruptive to users attempting to reach the application. These instances remind us that applications, much like the underlying networks they run on, can experience outages, and effective troubleshooting requires end-to-end visibility into both.
00:00 Welcome 00:14 Headlines: DNS and BGP and DDoS Attacks—Oh, My! This week we cover a couple of recent service degradation incidents involving DNS providers 2:19 Under the Hood: Kemal Sanjta, ThousandEyes’ resident BGP expert, joins us to discuss the May 6th disruption to Neustar’s UltraDNS service, which lasted nearly four hours. We discuss the BGP routing changes we observed during the incident and what they can tell us about the cause of the disruption. We also cover a separate incident involving Quad 9, a public recursive resolver service, which the company says was caused by a DDoS attack on May 3rd. 16:19 Expert Spotlight: Michael Batchelder (A.K.A., Binky), is here to discuss the two “Ds” of the Internet — DDoS attacks and the DNS Questions for Binky? Contact him at binky@thousandeyes.com 31:49 Outro: Questions? Feedback? Have an idea for a guest? Send us an email at internetreport@thousandeyes.com
This is The Internet Report, where we uncover what’s working and what’s breaking on the Internet—and why. Today, we focused on an interesting outage that impacted Cloudflare Magic Transit, a relatively new offering from the CDN provider which aims to efficiently route and protect the network traffic of its customers. On May 3rd at approximately 3:00 PM PDT (10:00 PM UTC), ThousandEyes vantage points connecting to sites using Magic Transit began to detect significant packet loss at Cloudflare’s network edge—with the loss continuing at varying levels, for approximately 2 hours.
While the outage impacted some Magic Transit customers more significantly than others, we also observed mitigation actions by at least one customer to avoid the outage and restore the availability of their service to their users. This outage reminds us that no provider is immune to outages, even cloud and global CDN providers. However, with proactive visibility, you can respond quickly to reduce outage impact on your users. Watch this week’s episode to hear more about the outage from the ThousandEyes perspective.
This is The Internet Report, where we uncover what’s working and what’s breaking on the Internet—and why. We’re joined this week by Hans Ashlock, Director of Technology & Innovation at ThousandEyes, to discuss Tuesday’s Microsoft Teams outage. On Tuesday, April 27th, ThousandEyes tests began to detect an outage affecting the Teams service starting around 3 AM (PT) and lasting approximately 1.5 hours. While the outage occurred in the overnight hours for much of the Americas, the global nature of the outage resulted in service disruption for users connecting from Asia and Europe.
Transaction views within the ThousandEyes platform show that Microsoft’s authentication service appeared to be available, however, the Teams application was unable to initialize, resulting in error responses. Watch this week’s episode to hear more about what ThousandEyes revealed about the nature of this outage—and what we can all learn from the incident.
This is The Internet Report, where we uncover what’s working and what’s breaking on the Internet—and why. On today’s episode, we’re thrilled to be joined by Kemal Sanjta, ThousandEyes’ resident expert on BGP. This week, we’re going under the hood on the April 16th BGP leak at Vodafone India, which leaked more than 30,000 prefixes, causing a major disruption of Internet traffic to some services. While some news outlets reported that the incident lasted approximately 10 minutes (starting around 1:50AM UTC or 9:50AM ET), we found that it lasted quite a bit longer—more than an hour in the case of some prefixes. Watch this week’s show to see how it impacted a major CDN provider.
This is The Internet Report, where we uncover what’s working and what’s breaking on the Internet—and why. We’re back from a short sabbatical to cover an interesting outage at Facebook in what appears to be an application outage compounded by a series of routing issues. On April 8th, for roughly 40 minutes, the Facebook application became unavailable for users around the globe who were attempting to connect to the service. Despite the short-lived nature of the outage, we observed prolonged performance degradation even after the application came back online for users. Suboptimal page load and response times, both of which can impact the user experience, were observed alongside a series of routing changes. This outage reminds us all of the importance of having visibility across network and application layers when troubleshooting and prioritizing issues that are impacting user experience. Catch this week’s episode to hear about the outage from ThousandEyes perspective.
On today’s episode, we discuss the recent outage on Verizon’s network that had widespread impacts on users in the US. ThousandEyes Broadband Agents detected an outage starting around 11:30am EST that manifested as packet loss across multiple locations concentrated along Verizon backbone in the US east coast and midwest. While the outage was resolved approximately an hour later, users connecting from the Verizon network across the US experienced varying degrees of impact, depending on the services they were connecting to. This serves as yet another reminder that the context around an outage directly affects the scope of the disruption. Watch this week’s episode to see what this outage looked like from ThousandEyes vantage points.
This is The Internet Report, where we uncover what’s working and what’s breaking on the Internet—and why. Despite a quiet last couple of weeks on the Internet, we started off our new year with quite the bang. As droves of mildly-caffeinated workers returned to their home offices on Monday after the holiday break, many were surprised to find that Slack was not available. On today’s episode, we go under the hood of Slack’s Monday outage to see what went wrong and how it was resolved. We’re also excited to be joined by Forrest Brazeal, a cloud architect, writer, speaker and cartoonist, to talk about everyone’s favorite subject: cloud resiliency. Watch this week’s episode to see the interview and hear our outage analysis. Show links: https://forrestbrazeal.com https://acloudguru.com https://cloudirregular.substack.com https://cloudirregular.substack.com/p/the-cold-reality-of-the-kinesis-incident
In this week's episode of #TheInternetReport... 00:00 Welcome 00:16 Headlines: About Monday’s Google Outage; Plus, Talking Holiday Internet Traffic Trends with Fastly 00:43 Under the Hood: This week, we go under the hood on a recent outage that took down the availability of several Google applications, including YouTube, Gmail and Google Calendar. Yesterday morning at approximately 6:50 AM EST, users around the world were unable to access several Google services for a span of around 40 minutes. While short-lived, the outage was notable in that it occurred during business hours in Europe and toward the beginning of the school day on the US east coast—so, people noticed, to put it bluntly. Catch this week’s episode to hear about the official RCA and what we saw from a network perspective. 10:18 Expert Spotlight: We’re thrilled to be joined by David Belson Senior Director of Data Insights, at Fastly talk about Internet traffic trends related to holiday online shopping and charitable giving. Cyber Five: what we saw during ecommerce's big week- https://www.fastly.com/blog/cyber-five-what-we-saw-during-ecommerces-big-week Decoding the digital divide- https://www.fastly.com/blog/digital-divide 19:14 Outro: We're taking a break for the rest of 2020 but join us on Jan. 05 2021 when we kick off the New Year with Forrest Brazeal: https://forrestbrazeal.com https://cloudirregular.substack.com
If you’re an AWS customer or rely on services that use AWS, you might have noticed the major, hours-long outage last week. On November 25th, at approximately 5:15 am PST, users of Kinesis, a real-time processor of streaming data, began to experience service interruptions. The issue was not network-related, and AWS later issued a detailed incident post-mortem analysis identifying an existing operating system configuration issue that was triggered by a maintenance event that involved adding server capacity. Over the course of the day, Amazon attempted several mitigation measures, but the outage was not completely resolved until approximately 10:23 pm PST.
What was notable about this outage was its blast radius, which extended far beyond AWS’s direct customers. Several AWS services that use Kinesis, including Cognito and CloudWatch, were affected, as were any user of applications consuming those services (e.g., Ring, iRobot, Adobe). This is a good reminder of the risk of hidden service dependencies, as well as the need for visibility to understand and communicate with customers when something’s gone wrong.
This is The Internet Report, where we uncover what’s working and what’s breaking on the Internet—and why. This week, we’re pleasantly surprised to say that the network did not break, and there were no major election-night outages to report. However, that’s not to say we didn’t catch performance glitches in the days and weeks around the big night. Watch this week’s episode, as we cover performance issues at a Secretary of State website as well as why CNN’s election map website was so slow to load for many.
We’ve got an election coming up here in the US, and over the last several weeks, we have been analyzing a dozen or so state election websites to take a closer look at how they’re hosted (e.g., do they use a CDN or are they self-hosted?) and to monitor them for outages. In this episode, we discuss the pros and cons of each hosting method and dive into some examples we’ve seen where election websites have had unexpected performance degradation. Catch this week’s episode to go under the hood on the websites powering the upcoming presidential election—and don’t forget to get out there and vote!
. In this week’s episode, we discuss two notable outages that happened last week. The first, at Twitter, took place on October 15 around 5:30 pm PST and impacted users’ ability to tweet or re-tweet. According to Twitter’s official statement, an internal system error was the culprit—putting to bed any theories of another hack. The second outage took place at the transit provider, Zayo, in the early morning hours of October 13. Although the outage seemed to mostly involve interfaces on the US west coast, Denver and the southwest (as well as a handful of other global locations), the impact of the outage was not very severe due to the time of the outage, which was outside of US business hours. Watch this week’s episode to hear more about these two outages.
This is The Internet Report, where we uncover what’s working and what’s breaking on the Internet—and why. In this week’s episode, we dive into a recent outage at Slack that caused intermittent issues for its enterprise users (including ourselves) for nearly a full day. The cause, as noted by Slack, was on the backend and related to an overloaded database. Next, we dig into another outage at Microsoft. According to their statement, a bug in an internal update seems to have revoked the routes to a number of devices that were believed to be unhealthy—thereby creating congestion in the rest of their network. This explanation jives with the increased packet loss we observed during this time period. Don’t miss this week’s episode, where we walk through these outages in depth
This is The Internet Report, where we uncover what’s working and what’s breaking on the Internet—and why. On today’s episode, we dive into a recent Azure AD disruption that significantly impacted access to Microsoft cloud services and apps (as well as third-party apps) for nearly three hours. We then went under the hood on a recent BGP hijacking in which Telstra began announcing routes to services that didn’t belong to it, such as Quad9. Catch this episode to hear our take on these incidents, and see below for show links, some additional commentary on these outages, and a sneak preview of next week’s episode.
On today’s episode, Angelique and I cover off on a couple outages that occurred over the past week. First, we discuss an application outage at Instagram that occurred on September 17th and lasted around 30 minutes. We also discuss a network outage on September 14th on the AWS backbone near Columbus, Ohio. This outage was a little more widespread, affecting nearly 100 interfaces and lasting around 30 minutes. Next, we dive into the upcoming bans on WeChat and TikTok, which have now been temporarily extended by a Federal judge, and then we walk through some of the network architecture differences between these two applications and how a potential shutdown could be enforced.
It was another quiet week on the Internet, so we wanted to spend some time answering your questions around some recent outages. Catch this episode as we discuss how you can understand the upstream relationships of the services you rely on to assess your risk profile. We also cover why SLAs fall short in protecting your business in the event of an outage, and why you need to proactively collaborate with your providers to solve issues faster.
The Internet held up reasonably well over the past week, all things considered. There were no major outages to report, which is a welcome repose for those impacted by the major outages the week prior. While it’s not an outage that occurred this past week, we did want to spend some time covering the recent Verizon Edgecast outage that occurred on August 21st. Watch this episode as we dive into this application-level outage to understand exactly what happened and who might have been impacted.
This is the Internet Report, where we uncover what’s working and what’s breaking on the Internet—and why. It was a rough week on the Internet last week, with outages and incidents across multiple services and providers including Slack, Zoom, AWS, and Verizon. However, in today’s episode we’re going to focus exclusively on Sunday’s CenturyLink / Level 3 outage that according to Cloudflare, caused a significant 3.5% drop in global Internet traffic, making it one of the most significant internet outages ever recorded.
his is the Internet Report, where we uncover what’s working and what’s breaking on the Internet—and why. On this week’s episode, Archana and I cover some recent outages that made headlines. This includes the Spotify outage, caused by an expired TLS certificate, that prevented users from accessing its platform. We also cover off on a widespread outage at Cogent during (what seems to be) a maintenance window. Then, we go “under the hood” on the prolonged outage at an IXP on August 18th to understand exactly what infrastructure was impacted and which downstream providers were subsequently impacted. We’re also joined by our guest, Prabhnit Singh, who currently leads ThousandEyes’ Internet & WAN product line, to discuss why we’re seeing an increased number of outages caused by expired TLS certificates and to cover some examples of past high-profile outages.
On this week’s episode, Archana and I cover recent headlines concerning social media platform, TikTok, and the gaming provider, Epic Games. TikTok appears to have gained some additional time (now 90 days) before the US government will enforce its ban on the service. Gaming provider, Epic Games, recently made news when its game Fortnite was removed from Apple’s App Store and Google’s Play Store for violating their Terms of Service. Epic was quick to file a lawsuit claiming the tech giants were in violation of anti-competition laws. The outcome of this case will be one to watch, and can have far-reaching impacts for developers. Next up, we speak with William Collins, Lead Cloud Architect at a Fortune 100 company, about cloud connectivity, on-ramp services and the difference between the “Big 3” on-ramp services.
This is the Internet Report, where we uncover what’s working and what’s breaking on the Internet—and why. On this week’s episode, Mike sat down with our guest, Ray Hunter, the senior network consultant at Globis in the Netherlands, to talk about SatComms and the role they play in connecting users, and what effect the mass deployment of Low Earth Orbital (LEO) satellites will have on networks and service delivery. We also discuss a recent move by the US to ban financial transactions between TikTok’s parent company, ByteDance, and US citizens, effectively removing financial incentives to serve US citizens. While not an outright ban, it does raise questions about how an outright ban even be enforced, and what that means for the broader conversation around Internet sovereignty.
This is the Internet Report, where we uncover what’s working and what’s breaking on the Internet—and why. On this week’s episode, Archana and I discuss a small number of outages that hit certain regions of the globe over the past week. This includes an outage that caused a midday disruption for people trying to connect to Reddit, a weekend DNS issue at Telstra, and a Cogent outage in EMEA and NA that had the signatures of a maintenance window. We also revisit Cloudflare’s root cause analysis concerning their recent DNS outage and answer some of the open-ended questions we had.
On this week’s episode, I am joined by Deepak Ravi from our Dublin technical sales engineering team to discuss a recent outage at Garmin. Garmin confirmed that it was a victim of a ransomware attack, which took down several of its services including its website functions, customer support, customer facing applications, and company communications. In this episode, we walk through what we observed in the ThousandEyes platform during the time of the attack, and what the impacts were on users attempting to access Garmin services.We’re also joined by ThousandEyes’ CISO, Alexander Anoufriev, to talk about what ransomware attacks are, how they manifest and how organizations can protect themselves against future attacks.
On this week’s episode, we cover a couple of significant application-layer outages at Github and WhatsApp that occurred over the past week. Then, Archana and I do a deep-dive into a network-related outage at Cloudflare that affected the availability of its popular DNS service for approximately 30 minutes. We’ll share what we saw through our vantage points in the ThousandEyes platform, and you can read Cloudflare’s full explanation of the incident on their blog/
On this week’s episode, we cover a recent move by the government of India to ban many Chinese-owned applications, including TikTok, which reportedly has more than 600,000,000 downloads in India. We also talk through a two-hour-long outage at Google Cloud Platform that affected multiple of its availability zones within a single region—highlighting that availability zones may be architected differently between providers—and briefly cover outages at Slack and Comcast, too. After our review of this week’s highlights, I sat down with Atif Khan, CTO of Alkira and former co-founder of Viptela to talk enterprise cloud strategy.
This week’s episode is brought to you by the letter “O” for outages — in particular, there were a number of broadband providers, globally, that suffered localized outages this past week. After we run down our top headlines, including a satellite provider rolling out managed SD-WAN, we take a look at outages in Comcast and AT&T’s networks. Make sure you join us next week to hear from Atif Khan, CTO at Alkira, as we talk about multi-cloud networking.
This is the Internet Report, where we uncover what’s working and what’s breaking on the Internet—and why. On this week’s episode, we cover a widespread T-Mobile outage that took down its cellular network for several hours and elicited a rare condemnation from the FCC. The culprit, according to the carrier, was a fiber cut—highlighting the need for redundancy and resiliency in the nation’s cellular networks. We also cover an issue with What’s App’s privacy settings that sent users scrambling to Twitter, as well as a recent move by Russia to “un-ban” the messenger app, Telegram. Then, stay tuned as we go one-on-one with Jason Black, the Head of Global Network Infrastructure at Uber Technologies, to discuss how Uber approaches its cloud architecture.
On this week’s episode, we discuss a recent BGP-related outage at a major public cloud provider, as well as a recent announcement that Cogent Networks has rolled out RPKI in an effort to strengthen its BGP route security. We’re also joined by Kemal Sanjta, principal engineer on our customer success team and our resident expert on Internet routing and security, to chat about these events. Catch this week’s episode here to dive into BGP with us.
On this week’s episode of the Internet Report, I’m joined by my colleague, Michael Batchelder (aka Binky), to discuss a DNS-related service disruption that affected users trying to access Amazon.com. We also talk about a recently discovered DNS vulnerability that could leave DNS providers susceptible to DNS amplification DDoS attacks. If you’re curious about what went wrong with Amazon’s service last week and want to know more about the role of DNS and why it’s so important, don’t miss this episode.
Welcome back to the Internet Report! On this week’s episode, we cover our usual check-up of ISPs, cloud and collaboration app outages, and discuss several major middle-of-the-night outages that affected services from providers such as Google and Virgin Media. We’re also joined by TeleGeography’s Alan Mauldin to discuss submarine cables, terrestrial networks, international Internet infrastructure and more.
Never a dull minute on the Internet! In today’s episode, Archana and I dove into a YouTube service disruption and an (unrelated!) Google network issue in India. We also discussed Slack’s explanation of their service disruption last week, and even talked through a case out of France where an education site experienced performance issues in lock step with time-of-day usage
On this week’s episode of The Internet Report, Archana and I cover some newsworthy updates that we’ve seen over the past week. We discuss a notable Facebook SDK outage that had ripple effects on other popular services that leverage its log-in functionality, including Spotify and Tik Tok. We also discuss a blog from AWS sharing their thoughts on the JEDI contract.
We’re also joined by Arash Molavi, the lead Internet researcher here at ThousandEyes. Arash shares his insight into outages we’re seeing, discusses what constitutes an outage, and why loss, latency and jitter can impact end-user experience in various ways depending on the context. Last, we cover our usual availability check of ISP, public cloud, and collaboration app provider networks.
On this week’s episode of The Internet Report, Archana and I are thrilled to be joined by Martin Levy, who is a distinguished engineer at Cloudflare focused on BGP route security and expanding Cloudflare's global network footprint. Check out this week’s episode to hear his thoughts on BGP security, best practices such as using RPKI and some recent routing incidents.
After speaking with Levy and hearing his perspective, we jump over to a discussion around some notable outages this past week, particularly one at Virgin Media that affected connectivity for users in the UK for several hours. After going through these events, we cover off on our usual health check of ISP, cloud, and collaboration service providers.
In this week’s episode, Archana and I are joined by a special guest, Christian Koch, who is the head of product, cloud and ecosystem at PacketFabric. Listen in as we cover the current state of global Internet health and dive into network outage numbers across ISPs, public cloud and collaboration platforms. This week, we saw that the overall number of outages wasn’t particularly concerning, reflecting our “new normal,” but there were a few notable outages that had far-reaching impacts. In particular, an outage at Tata Communications and a fiber cut at Level 3/CenturyLink had significant end-user impacts, as did another outage that took down access to GitHub.
In this week’s episode, Archana and I welcomed David Belson (@dbelson) of the Internet Society. We got to discuss some rather good news -- overall outage events are down more than 40% globally, and more than 44% in the U.S. after a several-week-long spike in events. We very well may be looking at our ‘new normal’.
It was yet another eventful week on the Internet, folks. In this week’s episode of The Internet Report, Archana and I discuss the latest figures around global Internet performance, noting that, despite an elevation in outages last month, the Internet is holding up well. ISP outages declined slightly in the U.S. and globally last week, but that wasn’t the case for UCaaS providers, who had a particularly rough time last week, especially in the United States. There was also a fairly large BGP route hijack on April 1 courtesy of Russian ISP, Rostelecom — the same ISP responsible for a route hijacking incident back in 2017. The prefixes involved belonged to Amazon, Cloudflare, and other services, and impacted the reachability of sites like Yelp.com.
Over the past month, ThousandEyes has been flooded with questions about how the Internet is holding up given the extra strain it’s been under with the sudden influx of remote workers, remote schoolers, and overall increased use due to COVID-19 related self-isolating and shelter in place orders. We’ve put out blogs and have conducted executive, media and analyst briefings. Network World and the IDG family of publications have even started publishing our data on a weekly basis to keep its readers up to date, as things are changing so frequently.
Because of the continued interest in how the Internet is handling the current and, potentially, increasing traffic loads, we decided that now is the right time to kick off a show to answer this question each week. How is the Internet faring? What were some of the most interesting events we observed during the week? We're pleased to share the inaugural episode of The Internet Report.
Listen along and don’t forget to subscribe to our blog and our YouTube Channel to be the first to get these episodes moving forward. And feel free to leave a comment here, on YouTube, or on Twitter, tagging @ThousandEyes and using the hashtag #TheInternetReport. We hope you find this info useful, and we look forward to your feedback.
Show Links: