The Debrief by incident.io : Recent Episodes

incident.io

In The Debrief, you'll hear conversations with engineers, product managers and founders about the ins and outs of incident response. Whether you're looking for actionable advice from folks who have been there or relatable stories, you'll find it on the Debrief.

View Details

We’re running a short mini-series on The Debrief podcast called Beyond the code, where we interview our engineers about what it’s really like to build at ⁠incident.io⁠.

In this episode, Product Engineer Rory B. and CTO Pete discuss how we’re using Claude Code and Git Worktrees to allow engineers to build multiple features in parallel. You can read more on our blog.

View Details

We’re running a short mini-series on The Debrief podcast called Beyond the code, where we interview our engineers about what it’s really like to build at incident.io.

In this episode, we chat with Product Engineer Leo about how we’re using AI tools like Claude Code to ship more product, more quickly.

View Details

We’re running a short mini-series on The Debrief podcast called Beyond the code, where we interview our engineers about what it’s really like to build atincident.io. In this episode, we chat with Product Engineer Leo about her time building On-call, our favorite engineering tooling, and what makes our engineering culture as good as cinnamon buns.

View Details

We’re running a short mini-series on The Debrief podcast called Beyond the code, where we interview our engineers about what it’s really like to build at incident.io.

In this episode, Norberto Lopes and Rory Malcolm discuss Rory's journey as a product engineer at incident.io, focusing on his experiences in the AI team and the challenges of developing the AI investigations product. They explore the engineering culture at incident.io and the impact of AI on incident management. The discussion also touches on the future of incident management and the evolving role of AI in a tech environment.

View Details

We’re running a short mini-series on The Debrief podcast called Beyond the Code, where we interview our engineers about what it’s really like to build at incident.io.

In this episode, Alicia chats with Kelsey, a long-time product engineer on the Response team. They dive into Kelsey’s journey from lab automation to building core incident features, the magic (and madness) behind Scribe, and what it’s like working on a high-performing team. Expect stories about AI, workarounds, and why versioning folders manually isn’t a scalable solution.

View Details

Join us for a deep dive into how incident.io is leveraging AI to build an intelligent incident investigator. Our guests, Ed and Lawrence, share insights on building AI-powered investigations that help teams to leverage huge amounts of data and signals to respond faster and more effectively.

View Details

In this episode, Stephen, Pete and Chris take a look back at 2024 at incident.io — reflecting on the year’s personal milestones, company-wide changes, and how the product has evolved along the way.

And as is customary, there's plenty of the usual good-natured humor along the way too.

View Details

In this episode, hosts Norberto and Lawrence discuss the recent CrowdStrike incident that began on July 19th.

You won't find any backseat commentary on the technical specifics, but instead a deep dive into the things we care about incident.io, like communication, their over response and proactive problem-solving during crises.

View Details

This week, we're talking to Sabin Roman, engineering manager at Linear, to talk about processes that sit behind building their product.

We cover how they build teams around planned work, how their "goalie" role works to protect teams from unplanned work, the zero bugs policy they've introduced and how they ensure everyone at Linear sweats the details on their product.

View Details

This week we sit down with Hank Jacobs, Staff Site Reliability Engineer at Netflix to discuss their deployment of incident.io across their organization.

Among other things, we discuss how great UX has allowed them to roll out to hundreds of teams in months, how they have more entries in their Catalog than any other incident.io customer, and how their partnership with incident.io has been an overall game changer.

View Details

During a recent episode of The Debrief, we spoke with Jeff Forde, Architect on the Platform Engineering team at Collectors, about building an incident management program at various stages of growth.

In that episode, we called it growth from zero to one, one to two, and two to three.

But what happens once you’ve scaled beyond three and answers to question you may have become that much harder to find.

To get to the bottom of this, we chatted with Oliver Tappin, Director of SRE at Eagle Eye, about what to do once your company has reached a point where there’s no precedent or roadmap, and you can’t necessarily look to others for answers.

View Details

This week, we have a really fun conversation lined up.

For this episode, we chatted with Toby Jackson, Global SRE Team Lead at Future, about why it’s a bad idea to take a cookie-cutter approach to incident management or, put another way, why it’s not a good idea to treat all incidents alike.

In our conversation, we discuss what’s wrong with this approach, some situations where this might actually make sense, how psychological safety factors into this conversation, and a whole lot more.

View Details

This week, we're sharing an extra special episode.

It's no secret that the decision to buy or build isn't exactly a straightforward one. And the decision you make can be influenced by a ton of factors.

But the fact is that in some instances, buying can make more sense than building, and in others, building can make more sense than buying.

In this episode, you'll hear from John Paris, Principal Engineer at Skyscanner, to get the story behind their build versus buy journey.

Joining him as the host for this episode is none other than the CPO of incident.io, Chris Evans.

In their conversation, Chris and John discuss Skyscanner's setup before adopting incident.io, what life has been like after adopting the platform, and a whole lot more.

View Details

It’s fair to say that AI is here to stay.

So, as companies grapple with this reality, they’re putting their best foot forward to build AI features that really make a difference for their customers.

But should you be building these features if there’s no obvious fit in your product? And even if there is, are you making sure to stay true to your product principles?

The reality is that deciding to build AI into your product isn’t a decision you make on a whim.

There are tons of considerations around how to do it right—many of which we wrestled with ourselves when we were building our AI features just a few months ago.

So, in this episode of The Debrief, we sat down with our CTO, Pete Hamilton, and Product Manager, Ed Dean, to get some perspective on how we weighed the decision to build with AI and how we thought about principles along the way.

View Details

It’s no secret that teamwork is one of those things that, when done right, can make a world of a difference.

So sometimes, when responding to a particularly complicated incident, it can be best to bring a team together to figure out what’s going on and work towards a fix.

But it’s not enough to just jam a bunch of folks into a room and hope for the best. You need a framework in place to ensure that everyone stays focused, diagnoses the issue and resolves it as quickly as possible.

And for SRE, Dan Slimmon, clinical troubleshooting is just the framework to help with this.

In this episode, we chat with Dan about this approach to collaboration and why, he thinks, it can help teams resolve issues much faster.

In our conversation we discuss what the benefits of clinical troubleshooting are, why teams get tripped up on collaboration in the first place, what firefighting and incident response have in common and a lot more.

View Details

Whether you’re a seasoned vet when it comes to incident response, or just getting started out, it can be easy to fall into the trap of doing too much all at once.

And it just makes sense.

Incident response is one of those things that doesn’t have a single, perfect formula, so teams can be left doing a little bit of everything in an effort to get it right.

That said there are some fundamentals that, regardless of how mature your organization is, can be a great launching off point to better incident response.

And that’s exactly what we’ll be talking about in today’s episode of the Debrief.

This time around, we’re joined by Viktor Stanchev of Anchorage Digital, to chat about actionable advice for responding to incidents—from declaration to post-mortem. We cover what having a good incident response even means, why it’s important to declare incidents early, how to better communicate during incidents and a whole lot more.

If you’ve been looking for practical advice for running incidents from a veteran in the space, you’re in the right place.

View Details

In last week’s episode of The Debrief, we had on Colette Alexander, Director of Engineering at HashiCorp, to discuss some of the myths around incident response.

In that conversation, one of the myths we spoke about was the idea that asking “why” is better than asking “how.” And how, in reality, asking "how" allows you to focus more on the contributing factors that led to an incident happening, whereas “why” tends to single out a person, which can lead to a lot of blame.

For this episode, we’re diving a bit deeper into the reasons “how” is not only better for learning, it’s also better for the psychological safety of your team.

This time around, we’re joined by Dennis Henry who currently works on the Architecture team at Okta. Dennis is a big believer in psychological safety and learning from incidents, so he’s just the person to shed light on this fascinating topic.

View Details

What if we told you that everything you thought you knew about incident response was wrong.

Well, at least some of it.

That some of the things you’ve been doing for years might not actually be having the impact you thought they did. Or, even worse, that some of the assumptions you’ve been making have actually been having a negative impact on you, your team and your organization.

This week, we’re talking about myths around incident response. And who better to dispel some of these myths than Director of Engineering at HashiCorp, Colette Alexander.

We chat about myths around learning and process, why “why” is the wrong question to be asking after incidents, and why documenting risk doesn’t necessarily help you manage them.

View Details

Whether you’re a seasoned company with 10+ years of operations, or a startup that’s just getting off the ground, making sure you have a good culture of engineering is really important.

Not only will this have a significant impact on the folks on your team, it’ll make a big difference with hiring.

When everyone knows that your company is the place to be when it comes to culture, attracting really good talent becomes that much easier.

But I was curious, what do some of the folks at incident.io think about engineering culture in general and how to best build it? Better yet, what about the engineering culture at incident.io? What’s it like?

To answer all of these questions and more, I sat down with Lisa Karlin Curtis, Tech Lead, and Alicia Collymore, Engineering Manager, to get their perspectives on this incredibly important topic.

We chat about what “culture” even means, why diversity is important, how teams can make sure their engineers feel empowered to share their perspectives and a whole lot more.

View Details

Q1 2024 is officially behind us.

So we figured that it was a great time for a bit of reflection on the exciting start to the year. In this episode, we sit down with our founders, Stephen, Chris, and Pete, to get a bit of perspective on how the last three months played out.

We chat about On-call, our AI launch, and the hundreds of other features, bug fixes, and bits of polish and delight that we've shipped over the last 12 weeks.

We also chat about the state of the company as a whole, our growth, and ultimately what's on the horizon.

View Details

Today, good incident communication isn't a nice to have—it's an absolute must.

But where do you even start? To help answer that question, we sat down with the VP of Engineering at SumUp, Adrián Moreno Peña, to get his perspective on how organizations of all sizes can share stellar comms no matter the situation. We discuss:

  • What it means to communicate during incidents
  • Why Status Pages are critical in helping to build trust
  • How you can have good comms even without a lead
  • ...and much more

View Details

Recently, we introduced our very first VP of Engineering, Norberto Lopes, to incident.io. As with all of our new joiners, we thought it would be helpful for folks to get acquainted with who exactly he is!

So in this episode of The Debrief, we'll do exactly that.

We sat down with Norberto to ask about his background, what he was doing before incident.io, what motivated him to join the company, and a whole lot more.

If you wanted an opportunity to get to know our VP of Engineering a bit more behind the scenes, then this is the episode for you.

Read Norberto's blog post explaining why he joined incident.io: https://incident.io/blog/why-i-joined-incident-io

View Details

Today, incident management is a core part of organizations, both big and small.

But what if you don't have an established incident management program, where do you start? Or what if you already have a program, but you're looking to optimize it a bit? Where do you start in that case?

Consider another situation: What if you're an established organization with years of incident management experience—what are some things that you can do to take things to the next level?

To talk through all of these scenarios and more, we sat down with Jeff Forde, Architect on the Platform Engineering team at Collectors.

Jeff has been a part of organizations at each of these phases, playing a key role in developing the incident management programs those organizations have today.

If you're looking for actionable advice on how to level up your incident management, then this is the episode for you.

View Details

This is on-call as it should be.

The secret's out. The world can finally know.

incident.io On-call is here.

Naturally, a lot of you may be wondering: why and why now. So to help answer those questions, we sat down with Chris and Pete, two of our co-founders here at incident.io to get a bit of background on this project:

  • What exactly went into it?
  • What were we hoping to solve for?
  • How are we addressing the pain points around being on call?
  • And most importantly, how are we stacking up against the incumbents in our space?

This episode will not only get you excited about this huge week, it'll get you pumped for what's ahead for on-call.

Learn more about on-call here: https://incident.io/oncall

View Details

Noting follow-up actions is really important at the end of the incident response process. The problem is that it can be really easy to overlook certain actions or forget to do them entirely.

With Suggested Follow-ups, this is now a thing of the past.

In this episode, you'll hear from Rob, the project lead for our latest Suggested Follow-ups feature, to get a peek behind the curtain. You'll hear him chat about:

  • What went into building Suggested Follow-ups
  • What the project timelines were
  • What were some of the challenges the team faced
  • ...and a lot more

You can listen to our ⁠AI announcement episode here⁠. Here are the rest of our AI episodes:

  • Related Incidents
  • Suggested Summaries
  • Assistant

View Details

For a lot of teams, incident management can be a bit of a headache.

It's stressful. It's not optimized. The whole process can feel like it's being held together with tape. Worst of all? Responders are the ones feeling the brunt of it. But in reality, your customers are, too. Think about it:

  • Incidents are running longer than they should
  • Follow-up actions to prevent similar incidents in the future aren't being followed through with
  • Learning from incidents doesn't happen at all due to the lack of documentation
  • ..and the list does go on and on

But honestly, the situation doesn't even have to be so dire. Things can be, generally speaking, totally fine. But you recognize that there are some things that you can do to make incident response really shine at your organization.

So if you're finding yourself looking for a better way, we've got you covered.

In this episode, you'll hear from two folks who have years of combined experience responding to incidents: Kerim, a Senior Developer Advocate at HashiCorp, and Lawrence, a Product Engineer here at incident.io.

The topic of our conversation? How to make incidents less painful. They discuss:

  • The first incident they experienced that made them realize the value of a good incident response process
  • Why teams aren't prioritizing incident response
  • What the value of responding to incidents is
  • What a good incident response process looks like
  • ...and more

View Details

Imagine an AI assistant that could automatically surface a whole host of useful incident response data points with just a prompt. Well, you won't need to imagine for much longer. That's exactly what we built in Assistant, one of our newest features powered by AI.

In this episode, you'll hear from Charlie, the project lead for Assistant, to get a peek behind this game-changing product. You'll hear him chat about:

  • What went into building Assistant
  • What the project timelines were
  • What were some of the challenges the team faced
  • What was it like learning the ropes of prompt engineering while building out Assistant
  • ...and a lot more

You can listen to our ⁠AI announcement episode here⁠ and our previous ⁠episode on Suggested Summaries here⁠.

Read our blog about Assistant.

View Details

Incident summaries are the source of truth for responders joining an incident at any point. But the reality is that with so many things happening at once—like needing to respond to the actual incident—updating these summaries as things play out can fall by the wayside. Enter, Suggested Summaries, one of our newest features powered by AI.

In this episode, you'll hear from Milly, the project lead for Suggested Summaries, to get a peek behind the curtain. You'll hear her chat about:

  • What went into building Suggested Summaries
  • What the project timelines were
  • What were some of the challenges the team faced
  • What was it like learning the ropes of prompt engineering while building out Suggested Summaries
  • ...and a lot more

You can listen to our AI announcement episode here and our previous episode on Related Incidents here.

View Details

For financial services companies, good incident management is absolutely critical—maybe more so than in other industries. So, for Michael Cullum and his team at Bud Financial, the choice to build an incident response tool felt right for them in the moment.

But very quickly, Michael and the team came face-to-face with the myriad limitations that come with building your own response tooling. Soon after, they would adopt incident.io to replace that internal Slackbot and save money, resources, and headaches in the end.

The end result? A total transformation of the way they managed incidents.

In this conversation, Chris Evans, CPO of incident.io, sits down with Michael to chat through this journey. Read more about Michael's buying decision: https://incident.io/customers/bud

View Details

Recently we went live with one of our biggest product launches to date AI. And this product was unique in that it was broken up into four smaller projects:

  • Related incidents
  • Suggested summaries
  • Suggested follow-ups
  • Assistant

So naturally most folks might be wondering: What were the biggest differences between these projects and what went into actually building out each of these features?

In this episode, you'll hear from Rob and Isaac, both Product Engineers who played a really critical role in the building out of related incidents, to get a peek behind the curtain.

  • You'll hear them chat about:
  • What went into building related incidents
  • What their project timelines were
  • What were some of the challenges they faced
  • What was it like learning the ropes of prompt engineering while building out related incidents
  • ...and a lot more

As an aside, if you want to learn more about each of these features, we're releasing four mini-episodes diving into each of them.

View Details

This week was a particularly exciting one for us at incident.io. We launched not one, not two, but four AI-powered features to help folks get the most out of their incidents.

  • Assistant
  • Suggested summaries
  • Suggested follow-ups
  • Related incidents

In this episode of The Debrief, we sit down with Ed Dean, Product Analyst, and Charlie Revett, Product Engineer, to talk through all of these features and explain how they're already making an impact. We discuss:

  • How we made sure to balance autonomy with AI
  • Why we decided to build these features and why now
  • What it was like working with design partners throughout this process
  • What it was like learning prompt engineering while simultaneously building a product rooted in the practice
  • ...and much more

You can learn more about our AI features here: https://incident.io/ai

View Details

What a year 2023 was at incident.io! While it's hard to summarize 365 days, a few things stand out:

  • We launched a bunch of new products like Catalog and Status Pages.
  • We hired a ton and we're now sitting at nearly 80 employees as of December 2023.
  • We expanded into the U S opening up a brand new office just a few weeks ago.
  • ...and there's still so much more ahead of us

So as we close the curtain on 2023, we sat down with the three co-founders of incident.io to do a bit of reflection on the wild ride that was this year.

In this episode you'll hear them discuss challenges, big wins, moments of growth, what's next for us, and most importantly, what the three co-founders like most about one another.

Read our year-end blog post here: https://incident.io/blog/reflecting-on-a-momentous-2023

View Details

If you're on a data team, have you ever considered using an incident management tool to respond to pipeline issues? If the answer is no, then you might want to check out this episode.

Here, we chat with Jack, Data Analyst at incident.io, to better understand why data teams can—and should—look to incident management tools like incident.io to manage issues. We chat about:

  • Why there's a big push for data teams to adopt the practices of engineering teams—from tools to processes
  • Why incident management tools aren't just for engineers
  • How an incident management tool like incident.io can transform the way you respond to things like pipeline issues
  • ...and much more

Read Jack's blog post about incident management for data teams here: https://incident.io/blog/incident-management-for-data-teams

View Details

At incident.io, we like to push the pace to deliver feature requests and bug fixes quickly. What's one of the things that helps us move with pace? The Product Responder role.

In this episode, we sit down with Sam Starling, Product Engineer, to talk about this crucial role and how it, ultimately, gives us an upper hand. We chat about:

  • How the role came about
  • What responsibilities the role has
  • What Sam's week as a Product Responder looked like

...and more

Read more about the Product Responder here: https://incident.io/blog/how-we-leverage-our-product-responder-role

View Details

In this interview, we chat with Lisa Karlin Curtis, Tech Lead at incident.io, about running meetings that, well, don't suck. In it, she gives actionable advice for running your own meetings, emphasizes why empathy in the workplace is important, reflects back on bad meetings she's run, and more.

Read Lisa's blog post here: https://incident.io/blog/how-to-run-meetings-that-dont-suck

View Details

Earlier this year, we launched our Status Pages product. In this episode, we chat with one of the product engineers on that project, Dimitra, about how that launch went. We discuss:

  • What it was like working on such a big project
  • What the team learned
  • What it was like collaborating cross-functionally

...and more

Read more about Status Pages here: https://incident.io/status-pages

Read Dimitra's post here: https://incident.io/blog/achieving-pixel-perfect-polish

View Details

In this episode, we chat with Alicia Collymore, Engineering Manager at incident.io, about advice she'd give EM's to be successful in their roles.

View Details

In this week's episode, we're joined by Matt Huxtable, CTO at Ziglu (an e-money issuer, offering a variety of digital finance services, particularly well known for its cryptocurrency services).

Matt talks about how the engineering team at Ziglu has evolved over time, building an agile culture and why "keep it boring" is his mantra.

Chris, Pete and Matt cover how to context switch between solving and communicating during an incident, their most creative incident fixes and why AI isn't ready to solve incidents for us just yet.

In this episode we cover:[01:20] Introducing Matt and Ziglu

[05:45] The evolution of the team at Ziglu

[09:45] Horizontal career growth in Engineering

[12:15] Finding an incident solution at Ziglu

[18:00] Balancing fixing with communicating during an incident

[20:26] Could AI help you solve incidents?

[27:00] Matt’s most creative incident fix

For more information, head over toincident.io/blog/podcast

View Details

In this episode, Pete and Lisa discuss why great communication is essential to the success of any incident management process. From keeping your wider team in the loop to minimise disruption, to using customer communication to strengthen your brand when things go wrong, the team share their experiences and top tips for having a transparent incident communication culture.

In this episode we cover:[01:10] Why is communication so important?

[07:30] Practical advice on building a good internal communication culture

[15:20] The cultural challenges of incident transparency

[25:30] Building trust with customer communications

For more information, head over to https://incident.io/blog/podcast

View Details

In this podcast, the three incident.io co-founders Stephen, Chris and Pete take a trip down memory lane, revisiting the story of how they came to found incident.io and the major milestones of the first 12 months in business.

Key topics/timestamps:[01:36] The origins of incident.io

[07:00] Building Response at Monzo

[17:00] The side hustle begins

[22:00] The first demo

[27:00] The first off-site

[34:00] Moving into the old fire station

[41:00] Hiring a team

[47:00] What next?

For more information, head over to https://incident.io/blog/podcast

View Details

In this podcast, our panellists discuss the foundations that any team needs to put in place when designing their incident management process. Starting from the basics of defining what we really mean by an incident, to how to set your severity levels, roles and statuses, Chris and Pete share their tips for building solid foundations to run your incidents.  

In this episode, we cover:

(00:55) What is an incident?  

(06:35) Questions to ask to figure out whether or not to declare an incident 

(12:27) Can you declare too many incidents?  

(17:59) Defining your severities  

(23:34) Why you need incident statuses 

(31:15) Incident roles and responsibilities 

(36:29) Using structured data to learn from incidents  

For more information, head over to https://incident.io/blog/podcast

View Details

In this podcast, our panellists discuss what it means to build a successful on-call team. Drawing on their experiences at fast growing start-ups and scale-ups, incident.io co-founders Pete and Chris cover everything from who should be on the rota and how to build a compassionate on-call culture, to compensation structures and tips for operationalising on-call.

In this episode, we cover: (04:07) What is on-call and why is it important?

(06:25) Who should be in-call?

(09:13) Should all teams be responsible for their own on-call, or should there be a dedicated team?

(12:59) How can you build a compassionate on-call culture?

(17:23) Spotting (and stopping) on-call heroes

(27:58) On-call compensation and other incentives

(39:30) Tips for operationalising on-call

For more information, head over to incident.io/blog/podcast