Send us Fan Mail
How do you design, implement and evolve effective alerting of your services and systems?
This week I'm joined by Krisha Vinnakota, Senior SRE @ Microsoft to dive into this topic. We cover...
๐ฉโ๐ป Treating your alerts like production code
โ
The link between SLOs and alerting
๐ก๏ธ How do you decide what thresholds to set?
๐ Signal to noise ratio and how to manage it
๐ Chaos engineering and postmortems are levers to improve alerting
...and much more.
You can find Krishna on...
LinkedIn: https://www.linkedin.com/in/krishna-vinnakota-8a03408/
You can buy Slight Reliability merch here (Note: you cannot order the mugs outside of New Zealand):
https://slightreliability.digitees.co.nz/
You can find Stephen on:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Bluesky: https://bsky.app/profile/slightreliability.bsky.social
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
Send us Fan Mail
Why do some of the most mature organisations in the world still experience major incidents due to expired or incorrectly applied certificates?
This week I'm joined by SRE and DevOps expert Charlie Al-Batty to discuss this and many other things including...
โ๏ธ His open source project for native SLOs and SLIs in Azure
๐งBuilding tools and automation to solve problems
๐ซฉ Dealing with alert fatigue
๐พ AI enabling us to become super developers
๐ The role of certifications in tech
...and much more.
You can find Charlie on...
LinkedIn: https://www.linkedin.com/in/charlie-al-batty-1895b911a/
His open source Azure SLO & SLI project "Azure SLO Guardian" can be found on GitHub: https://github.com/IamCharli01/Azure-SLO-Guardian/tree/main
We also mentioned Ansible AWX on this episode: https://github.com/ansible/awx
You can buy Slight Reliability merch here (Note: you cannot order the mugs outside of New Zealand):
https://slightreliability.digitees.co.nz/
You can find Stephen on:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Bluesky: https://bsky.app/profile/slightreliability.bsky.social
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
Send us Fan Mail
Can you combine your career with personal adventure? What would it be like to live in a truck and travel the country while working remotely?
This week I'm joined again by SRE & DevOps legend Amin Astaneh in one of the most human interviews I've ever done. We explore...
๐ The impact of the post-Covid layoffs
๐ป The logistics of living in a mobile home
๐ฅ Staring a consultancy
โค๏ธ The impact of the nomad lifestyle on relationships
๐งธ Being woken up by a giant black bear
...and much more.
You can find Amin on...
LinkedIn: https://www.linkedin.com/in/aminastaneh/
His website: https://certomodo.io/
and the Reliability Rebels podcast: https://podcast.certomodo.io/
Companies that Amin partners with...
Nobl9 (SLOs as a service): https://www.nobl9.com/
LearnKube (Kubernes education): https://learnkube.com/
Amin also mentioned Go Fast Campers if you were interested (I have no partnership or even knowledge of this company, just putting it here in case anyone was interested in seeing what Amin's living space looks like): https://gofastcampers.com/
You can buy Slight Reliability merch here (Note: you cannot order the mugs outside of New Zealand):
https://slightreliability.digitees.co.nz/
You can find Stephen on:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Bluesky: https://bsky.app/profile/slightreliability.bsky.social
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
Send us Fan Mail
When you first start implementing SRE it's a good idea to find early wins. Implementing monitoring of the four golden signals + availability is something I'm experimenting with at the moment to give our SRE team momentum and to pave the way for SLOs and more advanced observability.
In this solo episode I share my experiences, including...
๐ What are the four golden signals?
๐ฆฅ Observability in an async processing context
๐ The power of tracking availability
๐คท Why SLOs can be challenging to kick off on day 1
๐ช An invitation to come on the podcast
...and much more.
You can buy Slight Reliability merch here (Note: you cannot order the mugs outside of New Zealand):
https://slightreliability.digitees.co.nz/
You can find Stephen on:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Bluesky: https://bsky.app/profile/slightreliability.bsky.social
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
Send us Fan Mail
As an engineer you get constant dopamine hits by solving technical problems. As a leader you're often working toward long term goals that span months or years. How do you stay motivated in that context?
This week I'm joined by technology leader and fellow Aucklander Cads Oakley. We cover...
๐ Don't chase the result, harvest the learning
๐ The payoff for being a leader
๐ Giving feedback and giving others permission to give it to you
๐ The role of dopamine
๐คท How much control do we really have in our work?
...and much more.
You can find Cads on...
LinkedIn: https://www.linkedin.com/in/cadsoakley/
Website: https://aurelia.nz/
Content mentioned during the episode...
Cads' article on the role of dopamine: https://www.linkedin.com/pulse/managing-dopamine-mind-cรฆdman-cads-oakley/
"Happy" by Derren Brown https://www.goodreads.com/book/show/30142270-happy
What Got You Here Won't Get You There by Marshall Goldsmith https://www.amazon.com.au/dp/1401301304
You can buy Slight Reliability merch here (Note: you cannot order the mugs outside of New Zealand):
https://slightreliability.digitees.co.nz/
You can find Stephen on:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Bluesky: https://bsky.app/profile/slightreliability.bsky.social
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
Send us Fan Mail
This week I repurpose a talk I just did at the JuniorDev meetup in Auckland. If you're new to SRE or observability then this is the talk for you. For the more seasoned listeners, it's a chance to see how my perspective and understanding has changed over the years.
In the episode I refer to the This Is Fine! podcast on resilience engineering: https://www.thisisfinepod.com/
You can read the Google SRE books for free online here: https://sre.google/books/
You can buy Slight Reliability merch here (Note: you cannot order the mugs outside of New Zealand):
https://slightreliability.digitees.co.nz/
You can find Stephen on:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Bluesky: https://bsky.app/profile/slightreliability.bsky.social
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
Send us Fan Mail
How do you ingest and store petabytes of telemetry every day in a cost effective and high performing way? How can you do this in a way which gives engineers the operational data they need to keep services running? How has this challenge be tackled in the past and what's been the evolution?
This week I'm joined by Observe co-founder Jacob Leverich to go deep into this topic. We discuss...
๐พ A deep-dive into the evolution of telemetry storage and where it's going
๐ฝ The advent of generic storage that handles metrics, logs, and traces well
๐ซถ Having empathy for people who need great observability but struggle to obtain it
๐คฒ Not holding telemetry data hostage
โ๏ธ Lessons from 8 years of running a start-up
...and much more.
You can find Jacob on...
LinkedIn: https://www.linkedin.com/in/jacob-leverich/
And find out more about Observe here:
https://www.observeinc.com/
In the episode Jacob referred to Google's Dremel analysis platform: https://research.google/pubs/dremel-interactive-analysis-of-web-scale-datasets-2/
And Apache Iceberg: https://iceberg.apache.org/
You can buy Slight Reliability merch here (Note: you cannot order the mugs outside of New Zealand):
https://slightreliability.digitees.co.nz/
You can find Stephen on:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Bluesky: https://bsky.app/profile/slightreliability.bsky.social
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
Send us Fan Mail
How do you take all the utopian ideas you read about in books and apply them to the reality of the organisations we work in?
This week I'm joined by leader, mentor, and coach Rob Roe to tackle this question. We discuss...
๐ช๏ธ The pitfalls of functional silos
๐คซ Is the annual budget a load of rubbish?
๐ How our management promotion systems are often broken
๐ซ The power of virtual teams
๐ Team interaction models
...and much more.
You can find Rob on...
LinkedIn: https://www.linkedin.com/in/robinsonroe/
References from the episode:
Rob's YouTube channel "Unlocking Change": https://www.youtube.com/@unlockingchange
Action Inquiry by Bill Torbert: https://gla.global/the-glp/action-inquiry/
Seven Transformations of Leadership: https://hbr.org/2005/04/seven-transformations-of-leadership
Trademe Team Topologies case study: https://teamtopologies.com/industry-examples/trade-me-journey-towards-a-thinnest-viable-platform
You can buy Slight Reliability merch here (Note: you cannot order the mugs outside of New Zealand):
https://slightreliability.digitees.co.nz/
You can find Stephen on:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Bluesky: https://bsky.app/profile/slightreliability.bsky.social
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
Send us Fan Mail
We spend a third of our life at work. It needs to be something we enjoy and something with purpose. Our work experience also impacts our family, friends, and our personal lives.
This week I'm joined by tech engineer, leader, and author Richard Bown to explore this and many other topics including...
๐ช๏ธ The difficulty in applying the ideas we read in books in real organisations
๐คซ When you want to implement a thing you can't talk directly about the thing
๐ Does change require senior leadership? Can it be grassroots?
๐ซ The importance of looking after yourself at work
๐ The experience of publishing a novel
...and much more.
You can find Richard's book "Human Software - A Life in I.T" here: https://humansoftwarebook.com/
...and you can find Richard on...
LinkedIn: https://www.linkedin.com/in/richard-bown/
References from the episode (excluding some of the TV shows we mentioned):
Rosegarden Linux music composition editing environment: https://www.rosegardenmusic.com/
Managing Humans by Michael Lopp: https://www.oreilly.com/library/view/managing-humans-biting/9781430243144/
Rands leadership Slack: https://randsinrepose.com/welcome-to-rands-leadership-slack/
The Goal by Eliyahu Goldratt: https://www.goodreads.com/book/show/113934.The_Goal
Local Hero (1983): https://www.imdb.com/title/tt0085859/
Turn the Ship Around! by David Marquet: https://davidmarquet.com/books/turn-the-ship-around-book/
Scrivener (book writing tool): https://scrivener.app/
You can buy Slight Reliability merch here (Note: you cannot order the mugs outside of New Zealand):
https://slightreliability.digitees.co.nz/
You can find Stephen on:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Bluesky: https://bsky.app/profile/slightreliability.bsky.social
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
Send us Fan Mail
When you become a people leader there is no manual. How can we not only learn leadership skills but practice them and build leadership muscle?
This week I'm joined by Orion Group Limited co-founder Xiao Zhang to discuss...
๐ The challenge of transitioning into people leadership
๐ช How we don't get fit by watching other people work out
โ Pausing as an act of active leadership
๐ The power of slack time for creativity and systems thinking
๐ Going below the waterline
...and much more.
You can find Xiao on:
LinkedIn: https://www.linkedin.com/in/xiao-zhang-nz/
You can find Orion Group Limited here: https://www.oriongrouplimited.com/
Xiao mentioned Brenรฉ Brown's new book Strong Ground https://brenebrown.com/book/strong-ground/
I mentioned The Phoenix Project by Gene Kim, Kevin Behr and George Spafford https://itrevolution.com/product/the-phoenix-project/
I couldn't find the article on slack time but this is pretty good: https://buildrightside.com/problem-upgrade-chart
You can buy Slight Reliability merch here (Note: you cannot order the mugs outside of New Zealand):
https://slightreliability.digitees.co.nz/
You can find Stephen on:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Bluesky: https://bsky.app/profile/slightreliability.bsky.social
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
Send us Fan Mail
This week I kick off the 2026 season with some news and we explore how to prepare for a new role.
You can buy Slight Reliability merch here (Note: you cannot order the mugs outside of New Zealand):
https://slightreliability.digitees.co.nz/
You can find Stephen on:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Bluesky: https://bsky.app/profile/slightreliability.bsky.social
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
Send us Fan Mail
From the day we invented computers we've been struggling to keep applications running and delivering services to the business. Is this latest wave of AI helping or hurting us?
This week I'm joined by Causely founder Shmuel Kliger to dive into...
๐ The three waves of AI hype over the decades (the history of AI)
โ ๏ธ The dangers of over-promising and under-delivering what AI can do
๐ง What is causal reasoning?
๐ฑ Is AI replacing SREs?
๐ฎ AI as a way to allow humans to solve higher level problems
...and much more.
You can find Shmuel on:
LinkedIn: https://www.linkedin.com/in/shmuel-kliger-1a91963/
You can find Causely here: https://www.causely.ai/
Shmuel mentioned 'The Book of Why' by Judea Pearl and Dana McKenzie which can be found here: https://www.amazon.com.au/dp/046509760X
You can buy Slight Reliability merch here (Note: you cannot order the mugs outside of New Zealand):
https://slightreliability.digitees.co.nz/
You can find Stephen on:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Bluesky: https://bsky.app/profile/slightreliability.bsky.social
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
Send us Fan Mail
What is operational intelligence and how is it different from observability or BI?
This week I'm joined by SquaredUp's VP of Innovation Adam Kinniburgh to answer that question and many more including...
โ What is operational intelligence?
๐ Relating observability back to customer, business, or revenue
๐ The value of giving stakeholders confidence
๐ Who bridges the gap between tech and business or engineers and leadership?
๐ฆ Correlation VS causation and our innate desire to build connections
...and much more.
You can find Adam on:
LinkedIn: https://www.linkedin.com/in/adamkinniburgh/
...and the Operationally Intelligent podcast here:
https://squaredup.com/operationally-intelligent/
You can find the Slight Reliability merch store here:
https://slightreliability.digitees.co.nz/
You can find Stephen on:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Bluesky: https://bsky.app/profile/slightreliability.bsky.social
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
Send us a text
How does leading platform teams differ from leading product teams?
This week I'm joined by experienced technology leader Dinesh Sukhija to answer that question and many more including...
โ What is a platform team?
โฝ Coaching engineers to focus on outcomes
โ๏ธ Connecting platform initiatives to business goals
โ Identifying the limiters in your team
๐ค Spreading knowledge and avoiding single points of failure
...and much more.
You can find Dinesh on:
LinkedIn: https://www.linkedin.com/in/dinesh-sukhija/
You can find Stephen on:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Bluesky: https://bsky.app/profile/slightreliability.bsky.social
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
Send us a text
How has my first two years as a manager in tech been? What have I learned? What do I need to work on?
This week I share my experiences over the past couple of years. I cover:
๐ฅ My recent close call with burnout
๐ซถ How I attempted to build a team culture
๐ช The importance of tough conversations
๐ฅฑ How roles and responsibilities might be boring to think about but is critical
โ What's next?
...and much more.
You can find me on:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Bluesky: https://bsky.app/profile/slightreliability.bsky.social
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
Send us a text
How could AI help human beings negotiate the mountains of telemetry we collect to get simple and fast insight?
This week I'm joined by Ottermon AI CEO and founder Checo Pacheco about the lifecycle of observability coverage and tooling within organisations and how AI is helping to find signals amongst the noise and reduce cognitive load for SREs. We discuss...
๐ The need for a layer of logic on top of our telemetry data
๐ฒ The observability lifecycle of a DevOps team
๐ถ How most orgs have many observability tools, and how we might make that work
๐คฏ Reaching the limits of what humans can comprehend as a reason for AI
๐ How poor documentation may become AI's downfall in the future
...and much more.
You can find Checo on:
LinkedIn: https://www.linkedin.com/in/checopacheco/
You can find more about Ottermon AI on their website: https://www.ottermon.ai/ or on LinkedIn: https://www.linkedin.com/company/ottermon/
You can find Stephen on:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Bluesky: https://bsky.app/profile/slightreliability.bsky.social
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
Send us a text
What is chaos engineering and how is it being used in 2025?
This week I'm joined by Gremlin CEO and founder Kolton Andrus to discuss...
๐ช๏ธ What is chaos engineering and what is its origins?
๐ชด How has it evolved over the year?
๐ค The role of AI agents in SRE work
๐ฐ Justifying the value of chaos engineering
๐โโ๏ธโโก๏ธ How do I get started?
...and much more.
You can find Kolton on:
LinkedIn: https://www.linkedin.com/in/kolton-andrus-77315a2/
And you can find out more about Gremlin's new reliability intelligence platform here: https://www.gremlin.com/technologies/reliability-intelligence
You can find Stephen on:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Bluesky: https://bsky.app/profile/slightreliability.bsky.social
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
Send us a text
What are Team Topologies? How can they be used to deliver value simpler and more effectively (and in a more humane way)?
This week I'm joined by Luke McManus to discuss...
โฐ๏ธ What are the four team topologies?
๐ Can we have too much collaboration?
โ Team interaction models
๐ Cognitive load
๐โโ๏ธโโก๏ธ Value dynamics mapping
...and much more.
You can find Luke on:
LinkedIn: https://www.linkedin.com/in/luke-mcmanus-agile/
Check out the recently released second edition of the Team Topologies book by Matthew Skelton and Manual Pais here: https://itrevolution.com/product/team-topologies-second-edition/
Or Unbundling the Enterprise by Stephen Fishman and Matt McLarty here: https://itrevolution.com/product/unbundling-the-enterprise/
You can find Stephen on:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Bluesky: https://bsky.app/profile/slightreliability.bsky.social
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
Send us a text
How do you begin contributing to an open source project? What's it like? What do you get out of it?
This week I'm joined by Wendy Ha who shares her unique story of joining the Kubernetes project and becoming a contributor. We explore...
โฐ๏ธ What it's like working on one of the biggest open source projects in the world
๐ The benefits of contributing to open source
โ How much time and effort does it take?
๐ The unique challenges of contributing from APAC (and the need for more contributors in Australia and New Zealand)
๐โโ๏ธโโก๏ธ How to get started
...and much more.
You can find Wendy on:
LinkedIn: https://www.linkedin.com/in/wendyha-sut/
Ways you can get started contributing to Open Source:
CNCF from Zero to Merge Program: https://project.linuxfoundation.org/cncf-zero-to-merge-application
LFX Mentorship Program: https://mentorship.lfx.linuxfoundation.org/#projects_all
Outreachy Mentorship Program: https://www.outreachy.org/mentor/
Google Summer of Code: https://summerofcode.withgoogle.com/
Kubernetes Release Team Shadowing: https://github.com/kubernetes/sig-release/blob/master/release-team/README.md#release-team-shadow
You can find Stephen on:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Bluesky: https://bsky.app/profile/slightreliability.bsky.social
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
Send us a text
As an #SRE how do you influence senior leadership to get support and priority for the things you care about?
To answer this question I'm joined by Nora Jones, founder of Jeli and now Head of Pricing, Product Strategy and Growth at PagerDuty. Our conversation touches on...
๐ค How understanding needs to flow both ways (between engineers and leaders)
๐จ Reliability is as much an art as a science
๐ Using napkin math to start conversations
๐ง Understand the system (your org) before trying to change it
๐ฌ Using micro-interactions to gradually implement change
...and so much more.
You can find Nora on:
LinkedIn: https://www.linkedin.com/in/norajones1/
You can find more about PagerDuty here: https://www.pagerduty.com/nlp/trial-sign-up/
You can find Stephen on:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Bluesky: https://bsky.app/profile/slightreliability.bsky.social
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
Send us a text
This week I do a retrospective on the Slight Reliability podcast.
๐ How many people listen to it?
โค๏ธ How do I feel about the show?
๐ What's going well?
๐ชด What could be better?
โ What's next for the show?
If you want to check out the podcast that came before Slight Reliability, you can find Performance Time archived on YouTube here:
https://www.youtube.com/@performance-time
You can find Stephen on:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Bluesky: https://bsky.app/profile/slightreliability.bsky.social
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
Send us a text
Have you burned out at work? What was your experience? How did you work through it?
This week I'm joined by the incredible Colette Alexander to discuss what burnout is, what it means, and we both share our personal experiences burning out at work. We cover...
๐ฅ What is burnout?
โ Why does it happen?
๐ซ What are the symptoms?
๐ฅ Fight, flight, or freeze
๐งโ๐ Advice on how to recover
...and much more.
Resources from the show...
Why you're so angry at work (and what to do about it) by Natalie Rothfels https://www.lennysnewsletter.com/p/why-youre-so-angry-at-work
Burnout (book) by Amelia and Emily Nagoski https://www.burnoutbook.net/
How to do nothing (book) by Jenny Odell https://www.penguinrandomhouse.com/books/600671/how-to-do-nothing-by-jenny-odell/
You can find Colette on:
LinkedIn: https://www.linkedin.com/in/colette-alexander-4168267/
You can find the This Is Fine! podcast here: https://www.thisisfinepod.com/
You can find Stephen on:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Bluesky: https://bsky.app/profile/slightreliability.bsky.social
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
Send us a text
This week I'm joined by the wonderful Hanson Ho to discuss the unique challenges and opportunities in making our mobile apps observable! We cover...
๐ฑ The mobile/backend observability divide
โ๏ธ The challenge of distributed tracing on mobile apps
๐ The entire device runtime environment matters for your app
๐ค The quest for user-centric mobile observability
โ
Advice on how to get started with mobile observability
...and much more.
You can find Hanson on:
LinkedIn: https://www.linkedin.com/in/hanson-ho/
Bluesky: https://bsky.app/profile/bidetofevil.wtf
You can find out more about Embrace at https://embrace.io/
You can find Stephen on:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Bluesky: https://bsky.app/profile/slightreliability.bsky.social
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
Send us a text
This week on the I'm joined once more by SRE leader Michelle Casey who gives a broad and shallow introduction to resilience engineering. We cover...
๐๏ธโโ๏ธ Reliability VS Robustness VS Resilience
๐งฉ What is a complex system?
๐ข Safety one/safety two
๐ง Mental models
๐ฉ Human error
...and so much more.
Resources from this episode:
Four concepts for resilience (paper) by Dr. David Woods https://www.researchgate.net/publication/276139783_Four_concepts_for_resilience_and_the_implications_for_the_future_of_resilience_engineering
Building and revising adaptive capacity sharing for technical incident response (paper) by Dr Richard Cook and Dr Beth Long https://www.researchgate.net/publication/344259449_Building_and_revising_adaptive_capacity_sharing_for_technical_incident_response_A_case_of_resilience_engineering
Systems Thinking for Incident Analysis (talk) by Laura Nolan from LFI Conf 23 https://www.youtube.com/watch?v=-uXGg3g2yps
How Complex Systems Fail (website) by Dr. Richard Cook https://how.complexsystems.fail/
A Tale of Two Safeties (book) by Erik Hollnagel https://erikhollnagel.com/A Tale of Two Safeties.pdf
From Safety One to Safety Two (book) by Erik Hollnagel https://www.england.nhs.uk/signuptosafety/wp-content/uploads/sites/16/2015/10/safety-1-safety-2-whte-papr.pdf
Resilience: It's not you, it's the System (talk) by Dr Carl Horsley https://www.youtube.com/watch?v=ugC3GTKt23U
Above the line / Below the line (paper) by Dr Richard Cook (not original link) https://www.researchgate.net/figure/Above-the-Line-Below-the-Line-framework-adapted-with-permission-Cook-Woods-2016_fig3_333091997
How Your Systems Keep Running Day After Day (talk) by John Allspaw https://www.youtube.com/watch?v=xA5U85LSk0M
Behind Human Error (book) https://www.amazon.com.au/Behind-Human-Error-David-Woods/dp/0754678342
The Field Guide to Human Error Investigations (book) by Sydney Dekker https://www.humanfactors.lth.se/fileadmin/lusa/Sidney_Dekker/books/DekkersFieldGuide.pdf
The Howie Guide (paper) by Dr Laura Maguire, Nora Jones and Vanessa Granda https://howie-guide.pagerduty.com/
Resilience Engineering: Where do I start? (website) by Lorin Hochstein https://www.resilience-engineering-association.org/resources/where-do-i-start/
The STELLA report (paper) https://snafucatchers.github.io/
DORA Communtiy Discussion - Resilience Engineering (discussion) https://www.youtube.com/watch?v=g3cEJ7njJbc
This Is Fine! (podcast) by Colette Alexander and Clint Byrum https://www.thisisfinepod.com/the-pod
Send us a text
This week on the 100th episode I'm joined by DevOps and Resilience Engineering legend John Allspaw to talk about learning (especially from incidents). We discuss...
๐ Classroom VS situated learning
๐ค The myth of the perfect handover
ITIL as a coping strategy to try and make sense of the organic, wild, and messy
๐ฅ How you cannot incentivise to avoid incidents (it doesn't work that way)
โค๏ธโ๐ฉน You can't understand how something is broken unless you know how it's supposed to work in the first place
...and much more.
Resources from this episode:
Pre-Accident Investigations by Todd Conklin https://www.amazon.com.au/Pre-Accident-Investigations-Introduction-Organizational-Safety/dp/1409447820
Working at the Center of the Cyclone by Dr. Richard Cook https://itrevolution.com/articles/center-of-the-cyclone-dr-richard-cook/
To join the Resilience in Software Foundation head over to: https://resilienceinsoftware.org/
You can find John on:
Website: https://www.kitchensoap.com/
LinkedIn: https://www.linkedin.com/in/jallspaw/
You can find Adapative Capacity Labs here: https://www.adaptivecapacitylabs.com/
You can find Stephen on:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Bluesky: https://bsky.app/profile/slightreliability.bsky.social
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
Send us a text
This week I'm joined by SRE leader Trent Hornibrook who shares a story about how he improved on-call early in his career, and then we explore the broader theme of focusing on the things that matter in observability, incident response, on-call, and beyond. We discuss...
๐ Empowering engineers to implement change in your org
๐งโ๐ผ Focusing on what matters (customer & business > technology)
๐ Not just adding more monitoring as the output of each PIR
๐ How autonomy can lead to accountability
๐ณ How to influence change in an organisation
...and much more.
You can find Trent on:
LinkedIn: https://www.linkedin.com/in/trenthornibrook/
You can find Stephen on:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Bluesky: https://bsky.app/profile/slightreliability.bsky.social
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
Send us a text
This week I'm joined by SRE leader Andrew Hatch from Cisco ThousandEyes to talk about a dirty word in the resilience community... root cause. In this excellent conversation we explore...
๐ Is the root cause of every incident the big bang?
๐ฆ How the value of root cause degrades as complexity increases
๐ซฃ That if the culture is not blameless, people will hide things
๐ณ Alternative approaches to root cause analysis such as branching timelines
๐ Getting someone without skin in the game to facilitate your blameless post-mortems
...and much more.
You can find Andrew on:
LinkedIn: https://www.linkedin.com/in/hatchman76/
Check out Andrew's SREcon21 talk 'Learning from Complex Systems' which covers many of the topics introduced in this episode: https://www.youtube.com/watch?v=5pKGW61Ryvo
You can find Stephen on:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Bluesky: https://bsky.app/profile/slightreliability.bsky.social
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
Send us a text
This week I'm joined by David Dick from 2 Steps to (finally!) discuss synthetic monitoring. We cover...
๐ค What is synthetic monitoring?
๐ฆพ What are the benefits and drawbacks to using it?
โข๏ธ Non-web based synthetics (the tough stuff)
๐น Combining RUM and synthetics
๐ซข Does synthetics need an OTEL-like framework?
...and much more.
You can find David on:
LinkedIn: https://www.linkedin.com/in/david-dick/
You can find more about 2 Steps at https://2steps.io/#
You can find Stephen on:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Bluesky: https://bsky.app/profile/slightreliability.bsky.social
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
Send us a text
This week I'm joined by Cin7 Engineering Director Milan Brown to unpack the challenges of technology management and leadership. We discuss...
โ๏ธ Theory X vs Theory Y management
๐ฃ๏ธ Intention based leadership and communication
๐ข Conditions in an org for people to thrive
๐ตโ๐ซ How do you learn to manage and lead?
๐ซค Managing people when you're not an expert in what they do
...and much more.
Resources mentioned during the episode:
Turn The Ship Around! (book): https://davidmarquet.com/turn-the-ship-around-book/
Agile Conversations (book): https://itrevolution.com/product/agile-conversations/
Drive (book): https://www.danpink.com/books/drive/
Radical Candor (book): https://www.radicalcandor.com/the-book/
The Team Canvas (technique): https://theteamcanvas.com/
The Enginer/Manager Pendulum (article): https://charity.wtf/2017/05/11/the-engineer-manager-pendulum/
Retromat (tool for running retrospectives): https://retromat.org/
You can find Milan on:
LinkedIn: https://www.linkedin.com/in/milan-brown/
You can find Stephen on:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Bluesky: https://bsky.app/profile/slightreliability.bsky.social
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
Send us a text
This week Leon Adato and I break down the state of applying for roles in tech. We cover...
๐ What a resume or CV is and is not
๐ค Leveraging your connections rather than relying on applying cold
๐ช How most job descriptions are works of fiction
๐ฆพ White-fonting to game AI resume assessment
๐งช Experimental ways we could recruit
...and our pitch for Kubernetes the Rock Opera (and much more)
You can find Leon's job postings weekly on his website:
https://www.adatosystems.com/category/joblistings/
You can find Leon on:
LinkedIn: https://www.linkedin.com/in/leonadato/
Bluesky: https://bsky.app/profile/leonadato.bsky.social
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Bluesky: https://bsky.app/profile/slightreliability.bsky.social
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
Send us a text
This week Priyam Kumar shares his story of moving from a massive organisation to a startup and the challenges and growth that came from that. We discuss...
๐ช War stories and examples of production incidents
๐ฉน The "hacks" we build to keep things running (and how maybe that's just normal)
๐ Keeping it simple... YAGNI (You Ain't Gonna Need It!)
๐งฏ The perils of getting stuck in reactive mode
๐ Areas of of learning if you want to get into SRE
...and much much more.
You can find Priyam on:
LinkedIn: https://www.linkedin.com/in/priyam-kumar/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Bluesky: https://bsky.app/profile/slightreliability.bsky.social
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
Send us a text
This week Michelle Casey shares her insights as a 'head of' engineering manager in the SRE context. This was one of my favourite conversations on the podcast so far. We cover topics such as...
๐คท๐ฝ Why move into leadership?
๐๏ธ Learning from other leaders
๐ What is unique about SRE leadership?
๐ Women in engineering leadership
...and we go through some feedback I got as a leader recently.
Resources that Michelle mentions during the episode:
The Five Dysfunctions of a Team (book): https://www.tablegroup.com/topics-and-resources/teamwork-5-dysfunctions/
The Phoenix Project (novel): https://itrevolution.com/product/the-phoenix-project/
The Unicorn Project (novel): https://itrevolution.com/product/the-unicorn-project/
How Complex Systems Fail (website): https://how.complexsystems.fail/
How Your Systems Keep Running Day After Day (talk): https://www.youtube.com/watch?v=xA5U85LSk0M
The Curse of the Systems Thinker (article): https://blog.relyabilit.ie/the-curse-of-systems-thinkers/
Confessions of an SRE Manager (talk): https://www.usenix.org/conference/srecon23americas/presentation/hatch
Gender Decoder (website): https://gender-decoder.katmatfield.com/
You can find Michelle on:
LinkedIn: https://www.linkedin.com/in/michelle-casey-00b39837/
Steve Licks Instagram: https://www.instagram.com/tailsofstevielicks?igsh=MWFhenVzdzh6Zmtudw%3D%3D
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Bluesky: https://bsky.app/profile/slightreliability.bsky.social
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
Send us a text
This week Adam and I get philosophical about what constitutes maturity in the field of observability. We tackle questions such as...
๐ธ Does your org treat observability as a cost centre or a value add?
๐ฅ Are you using observability reactively to solve problems? Or proactively to build better products and services?
๐ค Is your observability connected to your users and business in a meaningful way?
๐ Is monitoring the social media sentiment of your product part of observability?
...and much more.
You can find Adam at:
LinkedIn: https://www.linkedin.com/in/adam-toth-innovateq/
InnovaTeQ website: https://innovateq.io/
I mentioned the 'This Is Fine!' podcast about resilience engineering. Find it on Spotify or at https://www.thisisfinepod.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Bluesky: https://bsky.app/profile/slightreliability.bsky.social
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
Send us a text
In this episode I explore the challenges of achieving unified observability when integrating with SaaS products and services. I cover:
๐ The new wave of mega-complex SaaS
โ๏ธ Challenges integrating SaaS with our observability pipelines
๐ฉโ๐ฆฏ How the lack of SaaS autonomy limits the effectiveness of OpenTelemetry
๐ฐ Paying twice to ingest, store, and search telemetry
๐ Monitoring and predicting SaaS observability costs
...and much more.
Shout out to Mark Chiavaroli (and apologies for mispronouncing your surname multiple times), Damian Sharrock, and Reece Hewitt for bouncing ideas on this topic.
The 'Is it observable?' series can be found here: https://isitobservable.io/
...and you can find Henrik on LinkedIn: https://www.linkedin.com/in/hrexed/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Bluesky: https://bsky.app/profile/slightreliability.bsky.social
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
Send us a text
This week I check in and give an update on work, life, and my attempts at bringing to life SRE practices in the world of non-production environment management.
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
This episode was sponsored by SquaredUp. SquaredUp combines all your data with awesome dashboards, analytics, health rollup, and notifications, into a unified observability portal. Using a data mesh architecture, SquaredUp is a beautifully simple way to get instant access to the insights that matter, whenever you need them. If you want to know more head over to https://squaredup.com/ to sign up for your free account.
This week I'm joined by Karanveer Anand, SRE Technical Program Manager at Google to discuss blameless post-mortems. We cover:
๐ฆ
The recent Crowdstrike outage and their public post-mortem
๐ When do we do a blameless post-mortem?
๐ How do we do a blameless post-mortem?
โ
How do we make sure action items are followed through?
๐ฐ The power of learning from post-mortems created by other teams and orgs
...and much more.
You can find Karanveer on LinkedIn: https://www.linkedin.com/in/karanveer/
You can find Crowdstrike's preliminary post incident report here: https://www.crowdstrike.com/blog/falcon-content-update-preliminary-post-incident-report/
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
This episode was sponsored by SquaredUp. SquaredUp combines all your data with awesome dashboards, analytics, health rollup, and notifications, into a unified observability portal. Using a data mesh architecture, SquaredUp is a beautifully simple way to get instant access to the insights that matter, whenever you need them. If you want to know more head over to https://squaredup.com/ to sign up for your free account.
This week Zach Michel from https://middleware.io/ and I discuss the state of OpenTelemetry and what it means to adopt it. We cover:
๐ฉ๏ธ Achieving observability in a SaaS world
๐ฅซ Context propagation - the magic sauce of OTEL
๐ช The telemetry gateway concept and leveraging the OTEL collector
๐ชต The state of OpenTelemetry logging
๐ซ Making use of the OpenTelemetry community
...and much more.
You can find Zach on LinkedIn: https://www.linkedin.com/in/zamichel/
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
For a list of ways to interact with the OpenTelemetry community go to:
https://opentelemetry.io/community/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
This episode was sponsored by SquaredUp. SquaredUp combines all your data with awesome dashboards, analytics, health rollup, and notifications, into a unified observability portal. Using a data mesh architecture, SquaredUp is a beautifully simple way to get instant access to the insights that matter, whenever you need them. If you want to know more head over to https://squaredup.com/ to sign up for your free account.
In Episode 80 Niall Murphy talked about the need for SREs to be better at articulating the value of our work. In this episode I'm joined by ex-Googler and Engineering Director (SRE) at Culture Amp Artem Yakimenko about how we might achieve this.
We discuss both quantifiable and qualitative approaches including leveraging the untapped data in support tickets, customer sentiment and rankings, the relationship between finance and performance, the link between user design and performance, and so much more.
Books mentioned in the episode:
100 Things Every Designer Needs to Know About People
By Susan Weinschenk
https://www.amazon.com.au/Things-Every-Designer-Needs-People/dp/0321767535
You can find Artem on LinkedIn: https://www.linkedin.com/in/temikus/
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
This episode was sponsored by SquaredUp. SquaredUp combines all your data with awesome dashboards, analytics, health rollup, and notifications, into a unified observability portal. Using a data mesh architecture, SquaredUp is a beautifully simple way to get instant access to the insights that matter, whenever you need them. If you want to know more head over to https://squaredup.com/ to sign up for your free account.
In the world of SRE we constantly talk about defining SLOs, but what about evolving them over time? This week I chat with SRE Tech Lead Dom Finn about just that. We cover the relationship between reliability and user analytics, latency classes as a way to speak SLOs with business stakeholders, the role of NFRs and how the thresholds differ from SLOs, and much more.
Books mentioned in the episode:
The Beginning of Infinity: Explanations That Transform the World
By David Deutch
https://www.amazon.com.au/Beginning-Infinity-Explanations-Transform-World/dp/0143121359
Turn The Ship Around!
By David Marquette
https://davidmarquet.com/turn-the-ship-around-book/
You can find Dom on LinkedIn: https://www.linkedin.com/in/dom-finn/
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
This episode was sponsored by SquaredUp. SquaredUp combines all your data with awesome dashboards, analytics, health rollup, and notifications, into a unified observability portal. Using a data mesh architecture, SquaredUp is a beautifully simple way to get instant access to the insights that matter, whenever you need them. If you want to know more head over to https://squaredup.com/ to sign up for your free account.
This week I talk about the impact of SaaS-first technology strategies on the work of an SRE. I pose questions about observability, ownership, on-call, and how much control we have over reliability.
You can find the Bleeding Tech blog on Medium: https://medium.com/@stownshend
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
This week I chat with Dan Slimmon about applying the approach doctors use to treat patient symptoms during incident response.
You can find Dan's blog at https://blog.danslimmon.com/ or connect with him on LinkedIn here: https://www.linkedin.com/in/danslimmon/
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
This episode was sponsored by SquaredUp. SquaredUp combines all your data with awesome dashboards, analytics, health rollup, and notifications, into a unified observability portal. Using a data mesh architecture, SquaredUp is a beautifully simple way to get instant access to the insights that matter, whenever you need them. If you want to know more head over to https://squaredup.com/ to sign up for your free account.
This week I hear about all things Kubernetes from Komodor CTO and co-founder Itiel Shwartz. We chat about the promise that was made when Kubernetes first entered the industry, the challenge of getting developers engaged and capable of working in Kubernetes, my hate/hate relationship with Helm but its important contribution to the Kubernetes project, Kubernetes observability, and so much more.
You can find the Kubernetes for Humans podcast here:
https://komodor.com/blog/the-kubernetes-for-humans-podcast/
Or find out more about Komodor here:
https://komodor.com/
Or find Itiel on LinkedIn: https://www.linkedin.com/in/itiel-shwartz-18542853/
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
This episode was sponsored by SquaredUp. SquaredUp combines all your data with awesome dashboards, analytics, health rollup, and notifications, into a unified observability portal. Using a data mesh architecture, SquaredUp is a beautifully simple way to get instant access to the insights that matter, whenever you need them. If you want to know more head over to https://squaredup.com/ to sign up for your free account.
This week I sit down and have a discussion with Amin Astaneh (from Certo Modo) about CI/CD. We cover the power of the standard change as a way to navigate ITIL while still implementing DevOps practices, what to monitor to make your CI/CD observable, single piece flow, testing in production, and so much more.
You can find Amin on his company website https://certomodo.io, LinkedIn: https://www.linkedin.com/in/aminastaneh/ and Twitter: https://twitter.com/aastaneh
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
This episode was sponsored by SquaredUp. SquaredUp combines all your data with awesome dashboards, analytics, health rollup, and notifications, into a unified observability portal. Using a data mesh architecture, SquaredUp is a beautifully simple way to get instant access to the insights that matter, whenever you need them. If you want to know more head over to https://squaredup.com/ to sign up for your free account.
"Environment issues are just incidents that happened to occur in a non-production environment"... so why do we treat them so differently?
In this first episode of the 2024 season I reflect on how we handle incidents in non-prod environments.
(Note: Had a few issues with noise suppression in OBS Studio cutting off the start of some words, will sort it for the next episode)
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
This week I speak with co-author of the original SRE book + the SRE workbook, and renowned speaker Niall Murphy.
We chat about the state of SRE in the current macro-economic climate and how we're not yet doing a very good job at articulating the value of SRE to leaders, the relationship that velocity and reliability have, the value of new features versus reliability improvements, and much more.
You can find Niall at:
LinkedIn: https://www.linkedin.com/in/niallm/
X: https://twitter.com/niallm
Website: https://relyabilit.ie/
(and his company Stanza: https://www.stanza.systems/)
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
X: https://twitter.com/the_kiwi_sre
Instagram: https://www.instagram.com/slight_reliability/
Paige Cruz (from Chronosphere) is back. This week we discuss sampling. What is sampling? Why do it? What kinds of sampling are there?
You can check out Chronosphere's cloud native observability platform here: https://chronosphere.io/
You can find Paige on:
LinkedIn: https://www.linkedin.com/in/paigerduty/
X: https://twitter.com/paigerduty
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
X: https://twitter.com/the_kiwi_sre
Instagram: https://www.instagram.com/slight_reliability/
This week Valeska Victoria returns to share some of her experiences working as an SRE at eBay.
We look at the cascading effect of production issues in complex integrated environments (how there's often no single root cause), developer literacy of how infrastructure works, the importance of ownership and accountability of reliability, and much more.
You can find Valeska on:
LinkedIn: https://www.linkedin.com/in/valeska-victoria/
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
X: https://twitter.com/the_kiwi_sre
Instagram: https://www.instagram.com/slight_reliability/
This week I chat with Ankit Jain from aviator.co about developer experience.
We define developer experience and developer productivity, and how this applies to SRE. We discuss the growing expectation on developers and how this leads to frustration and burnout. We also explore how to measure developer experience and how to start working to make improvements.
You can check out Aviator's developer experience platform here: https://www.aviator.co/
You can find Ankit on:
LinkedIn: https://www.linkedin.com/in/ankitjaindce/
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
X: https://twitter.com/the_kiwi_sre
Instagram: https://www.instagram.com/slight_reliability/
This week I had the privilege of interviewing Liz Fong-Jones from honeycomb.io about DevRel, Developer Advocacy, and how that applies to SRE.
We discuss the difference between Developer Relations (DevRel) and Developer Advocacy, how Liz got into advocacy, how DevRel helps companies and the community, and some tips on how to get traction with SRE practices in your organisation.
You can check out Honeycomb's observability platform here: https://www.honeycomb.io/
You can find Liz on:
LinkedIn: https://www.linkedin.com/in/efong/
Website: https://www.lizthegrey.com/ (all her social/links are here)
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
X: https://twitter.com/the_kiwi_sre
Instagram: https://www.instagram.com/slight_reliability/
This week I had the honour of chatting with Steve McGhee (former Google SRE, current Google Reliability Advocate, and co-author of Enterprise Roadmap to SRE).
We discuss the evolution of SRE from where it began at Google and how it is being adopted by enterprises around the world now (and why this is happening). We talk about getting leadership support and how we get reliability taken seriously, the lies we tell ourselves to justify incidents and issues, leveraging transformation projects to bring SRE to life, how SLOs can act as the fulcrum between dev and ops, the fallacy of the pyramid model of reliability... and so much more.
You can find Steve at on:
LinkedIn: https://www.linkedin.com/in/stevemcghee/
X: https://twitter.com/stevemcghee
You can find Steve's book "Enterprise Roadmap to SRE" here: https://sre.google/resources/practices-and-processes/enterprise-roadmap-to-sre/
Steve also mentions the book "A Seat at the Table": https://itrevolution.com/product/a-seat-at-the-table/
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
X: https://twitter.com/the_kiwi_sre
Instagram: https://www.instagram.com/slight_reliability/
This week on Slight Reliability Stephen discusses observability vendor lock-in. What is it? What does OpenTelemetry do to help? What areas are yet to be solved?
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
This week we sit down and talk about SLOs with CPO and co-founder of Nobl9 Brian Singer.
We talk about the importance of reviewing operational effectiveness, getting buy in from leadership, using SLOs to reduce noise, how to implement SLOs within different cultures and structures, the parallels between security and reliability... and much more.
You can check out Nobl9's reliability and SLO platform here: https://www.nobl9.com/
You can find Brian on LinkedIn: https://www.linkedin.com/in/briantsinger/
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
Instagram: https://www.instagram.com/slight_reliability/
This week Stephen chats with Valeska Victoria about her time working as an SRE at eBay.
Valeska shares her data driven approach to SRE, having a voice as a less experienced engineer, handling incidents under high pressure, leveraging large language models to rapidly find the information you need during an incident, and much more.
You can check out PromptOps here: https://www.promptops.com/
You can find Valeska on LinkedIn: https://www.linkedin.com/in/valeska-victoria/
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
Instagram: https://www.instagram.com/slight_reliability/
This week Stephen chats with Dr. Vlad Ukis about his journey discovering, and then implementing SRE practices at Siemens Healthineers (which led to him writing a book).
They discuss how the evolution of infrastructure necessitates a shift in how we operate, the power of selling SRE practices, the SRE infrastructure used to build SLOs and reliability capabilities, how he implemented SLOs, and much more.
You can find Vlad's book "Establishing SRE Foundations" here: https://www.amazon.com/Establishing-Foundations-Step-Step-Organizations/dp/0137424604
You can find Vlad on LinkedIn: https://www.linkedin.com/in/dr-vladyslav-ukis-5172ba32/
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
Instagram: https://www.instagram.com/slight_reliability/
Amin Astaneh (from Certo Modo) is back to discuss his experience working as a production engineer (SRE equivalent) at Meta.
Stephen and Amin discuss what it's like interviewing for big tech, "you build it, you own it", different SRE engagement models, SRE at different sizes of organisation, socialising your SRE success as a way to get traction, and so much more.
You can find Amin on his company website https://certomodo.io, LinkedIn: https://www.linkedin.com/in/aminastaneh/ and Twitter: https://twitter.com/aastaneh
The books Amin mentions are...
The Practice of Cloud System Administration: https://www.oreilly.com/library/view/practice-of-cloud/9780133478549/
Leading Change:
https://www.kotterinc.com/bookshelf/leading-change/
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
Instagram: https://www.instagram.com/slight_reliability/
This week Stephen talks to Praveen Kasam from Diconium Digital Solutions about how he led SRE transformations.
Praveen shares his experience transitioning from development to SRE and how leveraging automation and bringing application knowledge to the ops team provided quick wins. He also covers how he later applied SRE concepts to uplift the wider organisation. If you are out there looking for advice on how to implement SRE in your organisation, this is the episode for you.
You can find Praveen at:
LinkedIn: https://www.linkedin.com/in/kasampraveen/
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
X: https://twitter.com/the_kiwi_sre
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
This week Stephen asks Eric Schabell (Director of Technical Marketing & Evangelism @ Chronosphere) about how dashboards fit into modern observability.
They discuss how untamed observability can lead to unexpectedly high cloud bills, the similarities between dashboards and documentation, the "know > triage > understand" workflow, and much more.
You can find Eric at:
LinkedIn: https://www.linkedin.com/in/ericschabell/
X: https://twitter.com/ericschabell
And you can find Chronosphere at: https://www.linkedin.com/company/chronosphereio/
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
X: https://twitter.com/the_kiwi_sre
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
This week Stephen chats with Jamie Allen (Cheif Technologist AWS & SRE @ EPAM Systems) and Adam Kinniburgh (VP Innovation @ SquaredUp) about the concept of a single pane of glass (SPOG) for SRE.
Is it performance art or something actionable? Can alerting replace the need for dashboards? And are metrics drowning in the wake of distributed tracing?
You can find Jamie at:
LinkedIn: https://www.linkedin.com/in/jlallen/
And the Single Pain of Glass article he wrote here: https://medium.com/site-reliability-engineering-leadership/the-single-pain-of-glass-6e42930e966
You can find EPAM at https://www.epam.com/
And you can find the Google Dapper paper here: https://static.googleusercontent.com/media/research.google.com/en//archive/papers/dapper-2010-1.pdf
You can find Adam at:
LinkedIn: https://www.linkedin.com/in/adamkinniburgh/
X: https://twitter.com/adamkinniburgh
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
X: https://twitter.com/the_kiwi_sre
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
This week Stephen brings back Kyle Forster from RunWhen to talk about the purple elephant in the roomโฆ โAIโ.
What makes it GenAI, LLM, Advanced Statistics, or ML? Kyle shares his experience surrounding building AI powered search engines for SRE troubleshooting commands and how to incorporate a (paid) open source community of experts rather than trust AI by itself. They discuss what search looks like under the hood, why GenAI powered chatbots will or won't take over the SaaS industry, how Digital Assistants can be utilised by SREs to increase productivity (hint: giving them to app developers!), how to make informed decisions when purchasing AI products, and much more.
You can find Kyle at:
LinkedIn: https://www.linkedin.com/in/kyforster/recent-activity/all/
And you can find out more about RunWhen at:
Website: https://www.runwhen.com/
Product videos: https://www.youtube.com/@whatdoirunwhen
RunWhen Local: https://github.com/runwhen-contrib/runwhen-local (RunWhen Local is an open source troubleshooting cheat sheet that suggests commands from the RunWhen community for all of the namespaces in your cluster - ready to copy & paste)
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
X: https://twitter.com/the_kiwi_sre
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
This week Stephen chats with the internet incident librarian herself, Courtney Nash. They explore what Courtney has learned through meta-analysis of the over ten thousands incidents in the Verica Open Incident Database (VOID). They cover why MTTR needs to go in the garbage, joint cognitive systems, the value of looking at near misses and much more.
You can check out the VOID here: https://www.thevoid.community/
The two papers mentioned are:
Ironies of Automation by Lisanne Bainbridge: https://queue.acm.org/detail.cfm?id=3380779
Managing the Hidden Costs of Coordination by Laura Maguire: https://ckrybus.com/static/papers/Bainbridge_1983_Automatica.pdf
You can find Courtney at:
LinkedIn: https://www.linkedin.com/in/nashcourtney/
Twitter: https://twitter.com/courtneynash
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
This week Stephen chats with Martin Thwaites from Honeycomb about how developers can leverage observability to understand what they're building better, solve bugs quicker, and have more time for coding. They also discuss OpenTelemetry (the protocol and semantic conventions), manual versus automatic instrumentation, and how keeping every span of trace data is irresponsible.
You can find Martin at:
LinkedIn: https://www.linkedin.com/in/martin-thwaites-ab445120/
X: https://twitter.com/MartinDotNet
And Honeycomb at https://www.honeycomb.io/
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
X: https://twitter.com/the_kiwi_sre
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
Observability is a necessary adaptation to make sense of software systems in the Digital Age, but how can we unlock its power for non-engineer stakeholders (such as executives, product owners, etc)? Perhaps we need a layer of abstraction sitting on top of our detailed observability to get the most out of it.
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
This week Stephen chats with former-Google SRE Matt Brown about being on-call. They cover how to up-lift junior engineers so they can be on-call, what a fair on-call schedule looks like, run-books, and much more.
As you heard, Matt believes flexibility is key to a healthy on-call rotation. Matt is exploring ideas for improvements to existing tooling and products in this space and would love to hear from as many listeners as possible with feedback on what they find useful or frustrating with the existing tools they use to support on-call in their teams. You can reach him at oncall-feedback@mkmba.nz or schedule a chat via https://zcal.co/mattb/oncall, please don't be shy!
You can also find Matt at:
Website: https://www.mattb.nz/
LinkedIn: https://www.linkedin.com/in/mattbrown/
Mastodon: https://mastodon.nz/@mattb
Twitter: https://twitter.com/xleem
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
The internet is full of people who want to tell you about SRE, DevOps, and Platform Engineering and how different and similar they are... and will give you the impression that these things compete with each other. But do they? And is it a helpful question to ask in the first place?
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
In this episode Amin Astaneh from Certo Modo discusses his experience undertaking an SRE transformation over several years.
Stephen and Amin cover a lot of ground including making ops work visible, measuring toil, the power of calculating the $ value of work, getting developers on-call, the embedded model for SRE, SLOs, culture change, and a whole lot more.
You can find Amin on his company website https://certomodo.io, LinkedIn: https://www.linkedin.com/in/aminastaneh/ and Twitter: https://twitter.com/aastaneh
The books Amin mentions are...
The Practice of Cloud System Administration: https://www.oreilly.com/library/view/practice-of-cloud/9780133478549/
The Phoenix Project: https://www.oreilly.com/library/view/the-phoenix-project/9781457191350/
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
Instagram: https://www.instagram.com/slight_reliability/
In this episode Stephen Townshend and Sonja Chevre from Tyk discuss making APIs observable, and some anti-patterns to avoid. They cover GraphQL, OpenTelemetry and semantic conventions, correlation IDs, observability pipelines, and much more.
You can find Sonja on LinkedIn: https://www.linkedin.com/in/sonjachevre/ and Twitter: https://twitter.com/SonjaChevre
You can listen to Sonja's KubeCon talk here: https://youtu.be/IkEUJjRBCbo
You can find Tyk's open source gateway here: https://github.com/TykTechnologies/tyk
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
Instagram: https://www.instagram.com/slight_reliability/
In this episode Stephen Townshend and Harinder Seera explore how to monitor and manage the cost of cloud. They discuss FinOps as a cultural practice, anti-patterns for implementing in the cloud, keeping cost down through resources, pricing, and architecture... and much more.
You can find Harinder on LinkedIn: https://www.linkedin.com/in/harinderseera/
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
Instagram: https://www.instagram.com/slight_reliability/
In this episode Stephen shares his experiences traveling overseas to the UK and Singapore AWS Summit, SREcon APAC, and the internal SquaredUp conference "SqUpCon".
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Slight Reliability artwork on Instagram:
https://www.instagram.com/slight_reliability/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
Instagram: https://www.instagram.com/slight_reliability/
In this episode Stephen discusses the role of dashboards within the context of the Digital Era. What are they not appropriate for? What can they help with? What kinds of things are suitable to present?
If you want to get involved in the SquaredUp dashboard competition head along to: https://squaredup.com/blog/dashboard-competition/ (everyone who submits an entry gets a t-shirt, you can also win Star Wars Lego, get video interviewed by me, and have the story of your dashboard presented both as a blog and on our Dashboard Gallery).
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
Instagram: https://www.instagram.com/slight_reliability/
This week Bruce Cullen is back to share his experiences from KubeCon + CloudNativeCon 2023 Europe. We chat about OpenTelemetry, green engineering, securing your CI/CD pipeline and much more.
Bruce is the Director of Engineering at SquaredUp. You can find him on LinkedIn: https://www.linkedin.com/in/bruce-cullen/
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
If you like Slight Reliability's mspaint style artwork you can find more of it on Instagram: https://www.instagram.com/slight_reliability/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
In this episode Stephen Townshend chats to Andy Thurai (VP and Principal Analyst at Constellation Research) about Andy's latest report titled "Trends in Incident Management 2023". They chat about "mean time to innocence", status pages, they debate whether AI or ML has real value for incident management, and ponder why anyone would willingly decide to become an incident commander?
You can find Andy's report here: https://www.constellationr.com/research/2023-trends-incident-management
You can find Andy on LinkedIn here: https://www.linkedin.com/in/andythurai/
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
In this episode Stephen Townshend chats to Tim Wheeler (Director of Engineering Services at SquaredUp) about his work implementing and continually monitoring DORA metrics. They chat about customising each metric to your own unique context, avoiding the weaponisation metrics, the "tools will solve this for me" trap, and much more.
The books mentioned during this episode were: Accelerate, The DevOps Handbook, The Phoenix Project, The Unicorn Project, Lean Enterprise, and Sooner, Safer, Happier. Tim also mentioned the work of Bryan Finster (https://twitter.com/BryanFinster).
You can find Tim on LinkedIn: https://www.linkedin.com/in/timjameswheeler/
You can find out more about SquaredUp at https://squaredup.com/
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
In this episode Stephen explores the SRE concept of "toil". What is it? How can we measure it? How do we reduce it?
Also in this episode: Can we make non-technology systems observable? (like we do technology ones), and the ineffectiveness of change advisory boards (CAB). Also,ย Stephen's upcoming attendance at SREcon, AWS Summit, and SLOconf.
Shout outs to Steve McGhee, Dom Finn, and Shea Stewart.
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
In this episode Stephen Townshend and Anurag Gupta discuss the new reliability.org community for SREs or reliability engineers to share experiences, ask questions, and find community. They discuss the value of community and sharing your thoughts, collaboration between organisations, vicious versus virtuous cycles for reliability, and much more.
You can join us in the community by visiting https://www.reliability.org/
You can find Anurag:
On LinkedIn: https://www.linkedin.com/in/awgupta/
You can find out more about Shoreline by visiting https://www.shoreline.io/
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
In this episode Bruce Cullen interviews Stephen Townshend about the past, present, and future of the Slight Reliability podcast. They discuss their shared backgrounds in software testing, the different career paths that testing has opened up, and much more!
Bruce is the Director of Engineering at SquaredUp. You can find him on LinkedIn: https://www.linkedin.com/in/bruce-cullen/
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
In this episode Ivan Merrill from Fiberplane shares his experiences implementing observability within some of the large complex organisations he's worked for in the past.
You can find Ivan on LinkedIn: https://www.linkedin.com/in/ivan-merrill-1a05223/
You can find out more about Fiberplane here: https://fiberplane.com/
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
In this episode I discuss the word "insight" within the context of observability. Is insight something tools can provide? Is it something you can reproduce?ย
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
In this episode Stephen Townshend discusses our increased dependency on third party cloud services and what this means for reliability with Jeff Martens and Ryan Duffield from https://metrist.io/.
You can find Jeff...ย
On LinkedIn: https://www.linkedin.com/in/jmartens/
On Twitter: https://twitter.com/Jmartens
You can find Ryan...
On StackOverflow: https://stackoverflow.com/users/2696/ryan-duffield
On GitHub: https://github.com/rduffield
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
In this episode I propose the use of scatterplots of raw data to better understand how our systems are behaviour and what our customers are experiencing. The ideas from this episode come from my time as a performance engineer and working with legends in that space Richard Leeke (https://www.linkedin.com/in/richard-leeke-450448/) and Neil Davies (https://www.linkedin.com/in/neildaviesnz/).
For some basic examples of scatterplots and what they show you versus line charts check out an article I wrote back in 2017 called "Let's Talk About Averages": https://www.linkedin.com/pulse/lets-talk-averages-stephen-townshend/
Another proponent of scatterplots is Stijn Schepers (https://www.linkedin.com/in/stijnschepers/). Here's an article he wrote about it in 2019: https://www.linkedin.com/pulse/performance-testing-act-like-detective-use-raw-data-stijn-schepers/
Neil Davies' article on tornado scatters "Chasing Tornadoes" can be found here: http://www.performance-workshop.org/wp/wp-content/uploads/2013/12/Chasing_Tornadoes_Davies.pdf
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
In this episode we discuss uplifting telemetry knowledge within engineering teams to enrich their work (and their lives) with Paige Cruz from Chronosphere. We cover why not to take a chainsaw to your observability in order to cut costs, the dark side of auto-instrumentation, story telling with live data, and much more.
The book that Paige recommends at the end is "Effecting Monitoring and Alerting for Web Operations": https://www.oreilly.com/library/view/effective-monitoring-and/9781449333515/
You can check out Chronosphere here: https://chronosphere.io/
You can find Paige on LinkedIn: https://www.linkedin.com/in/paigerduty/
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
In this episode we discuss cognitive overload in SRE with Paige Cruz from Chronosphere. We cover both what cognitive load is, what causes it, as well as some potential antidotes and preventative measures.
You can check out Chronosphere here: https://chronosphere.io/
You can find Paige on LinkedIn: https://www.linkedin.com/in/paigerduty/
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
In this episode I discuss my "bigger picture" perspective of what observability needs to be, and why it's important we include business and customer into what we monitor in the Digital Era.
The books I highlight in this episode are...
Observability Engineering https://www.oreilly.com/library/view/observability-engineering/9781492076438/
Sooner, Safer, Happier: https://soonersaferhappier.com/book/
The Phoenix Project https://www.oreilly.com/library/view/the-phoenix-project/9781457191350/
The Unicorn Project https://www.oreilly.com/library/view/the-unicorn-project/9781098124175/
Accelerate: https://www.oreilly.com/library/view/accelerate/9781457191435/
You can grab a copy of the 2022 State of DevOps report at: https://cloud.google.com/devops/state-of-devops
The blog I mentioned was The Insight Industrial Complex: https://benn.substack.com/p/insight-industrial-complex
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
In this episode we speak to Josรฉย Velez from Rely about reliability at scale, a top down approach to SLOs, the potential and limitations of AI and ML in operations, the question of service ownership, utilising the business criticality of services in how we monitor the underlying infrastructure, and much more.
You can check out Rely at https://www.rely.io/
You can find Josรฉ on LinkedIn: https://www.linkedin.com/in/josevelez-relyio/
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
In this episode we speak to Ken Hamric about distributed tracing, leveraging tracing for better testing, and observability driven development.
The tool that Henrik Rexed integrated with Tracetest was Kuberhealthy (https://www.cncf.io/projects/kuberhealthy/) and you can watch a video of him discussing it in combination with Tracetest here: https://youtu.be/PKQQEeeMYxg?t=2492
Ken also mentioned Charity Majors' writing about observability driven development: https://thenewstack.io/a-next-step-beyond-test-driven-development
You can check out Tracetest:
- The official website: https://tracetest.io/
- GitHub repo: https://github.com/kubeshop/tracetest
- Discord channel: https://discord.com/channels/884464549347074049/963470167327772703
You can find Ken on LinkedIn: https://www.linkedin.com/in/ken-hamric-016b1420/
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
In this episode Stephen explores the pros and cons of centralising observability data. Is it a practical to stand up a complex and costly data storage and retrieval solution? Is there another way?
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
This week I am joined by Ana Margarita Medina and Adriana Villela, the hosts of the On-Call Me Maybe podcast, to discuss what we'd like to see for SRE in 2023. We talk about observability, SRE recruitment, what organisations need in place to set SRE up for success, and much more.
You can find the On-Call Me Maybe podcast on most podcast platforms or go directly to the website here: https://oncallmemaybe.com/
Twitter: https://twitter.com/oncallmemaybe
Mastodon: https://mastodon.social/@oncallmemaybe
You can find Adriana on:
LinkedIn: https://www.linkedin.com/in/adrianavillela/
Twitter: https://twitter.com/adrianamvillela
Mastodon: @adrianamvillela@hachyderm.io
Blog: https://adri-v.medium.com/
You can find Ana on:
LinkedIn: https://www.linkedin.com/in/anammedina/
Twitter: https://twitter.com/Ana_M_Medina
Mastodon: @anamedina@hachyderm.io
You can find me on:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
To begin 2023 I share the books I read last year in my quest to be a better SRE.
Here is a list of all the books mentioned during the episode:
You can find me on:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
This week Henrik Rexed and Stephen Townshend discuss their New Year's resolutions for observability. They cover OpenTelemetry and a unified query language, continuous profiling, raw data analysis, instrumenting code, using distributed tracing as part of testing, and much more.
Some of the tools or resources mentioned during the episode include:
https://tracetest.io/ (distributed tracing for testing)
https://github.com/open-telemetry/opamp-go (OTEL orchestration)
https://ebpf.io/ (for continuous profiling)
You can find Henrik on LinkedIn: https://www.linkedin.com/in/hrexed/ and Twitter: https://twitter.com/Hrexed
You can find the Is It Observable? series on YouTube: https://www.youtube.com/@IsitObservable
And the Perfbytes Podcast on most podcast platforms: https://www.perfbytes.com/p/perfbytes.html
You can find me on:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
This week we talk to Steve Gill and Gwen Berry from IAG to discuss their experiences forming an SRE incubator team (starting SRE from scratch in a large enterprise). We discuss on-call, SLOs, single pane of glass, pivoting, chaos engineering, and much more.
You can find Steve on LinkedIn: https://www.linkedin.com/in/stevegill239/
You can find Gwen on LinkedIn: https://www.linkedin.com/in/gwen-berry-56324418b/
You can find me on:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
This week I share the observations I made at AWS re:Invent relating to SRE work including the lack of SREs at the event, data warehouses for observability data, the use of topologies to understand complexity, FinOps, serverless, making sense of enormous amounts of data... and more.
You can find me on:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
This week I was at the AWS re:Invent conference in Las Vegas, so I took the opportunity to walk around the expo asking observability vendors what their perspective or definition of "observability" was (and reflected on that).
You can find me on:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
In this episode I explore the different kinds of SRE out there and the different needs they fill in the industry, and discuss some ethically dubious practices around hiring SREs.
You can find me on:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
Music from Uppbeat (free for Creators!).
Intro:
https://uppbeat.io/t/sensho/good-times
License code: QBXDSEGNJZY9DDIC
Outro:
https://uppbeat.io/t/mountaineer/voyager
License code: 5C0VMTUOULFSRSTM
In this episode I chat to Kyle Forster and Shea Stewart from RunWhen about the concept of "social reliability engineering" and how it could help SREs from organisations all over the world create an ecosystem of sharing and collaboration.
You can find Kyle on LinkedIn: https://www.linkedin.com/in/kyforster/
You can find Shea on LinkedIn: https://www.linkedin.com/in/sheastewart/
To find out more about RunWhen: https://www.runwhen.com/
And an example of the "street map view" of a tech stack: https://www.youtube.com/watch?v=SOvH9lcgCXg
You can find me on:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
Music from Uppbeat (free for Creators!).
Intro:
https://uppbeat.io/t/sensho/good-times
License code: QBXDSEGNJZY9DDIC
Outro:
https://uppbeat.io/t/mountaineer/voyager
License code: 5C0VMTUOULFSRSTM
In this episode I reflect back on the very first episode of Slight Reliability "What the heck is SRE anyway?" and see if my perspective has changed since then. I also tackle the confusion about what SRE is and is not.
Shout out to Sebastian Vietz (https://www.linkedin.com/in/sebastianvietz/) for his "Service Reliability Engineering" terminology and Richard Benwell (https://www.linkedin.com/in/richard-benwell-ab887b11/) for highlighting the way SRE offers a different value proposition depending on the scale of the services in question.
You can find me on:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
Music from Uppbeat (free for Creators!).
Intro:
https://uppbeat.io/t/sensho/good-times
License code: QBXDSEGNJZY9DDIC
Outro:
https://uppbeat.io/t/mountaineer/voyager
License code: 5C0VMTUOULFSRSTM
In this episode I announce my new role as Developer Advocate (SRE) at SquaredUp, and what this means for the Slight Reliability podcast.
You can find me on:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
Music from Uppbeat (free for Creators!).
Intro:
https://uppbeat.io/t/sensho/good-times
License code: QBXDSEGNJZY9DDIC
Outro:
https://uppbeat.io/t/mountaineer/voyager
License code: 5C0VMTUOULFSRSTM
In this episode I give a summary of the book Team Topologies by Matthew Skelton and Manual Pais (https://teamtopologies.com/book) and how this relates to implementing SRE practices.
(POINT OF CORRECTION: One of the authors is "Matthew" Skelton, not "Michael")
You can find me on:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
Music from Uppbeat (free for Creators!).
Intro:
https://uppbeat.io/t/sensho/good-times
License code: QBXDSEGNJZY9DDIC
Outro:
https://uppbeat.io/t/mountaineer/voyager
License code: 5C0VMTUOULFSRSTM
In this episode I give my take on the Accelerate State of DevOps 2022 from the SRE perspective.
You can find the Accelerate State of DevOps Report 2022 here: https://cloud.google.com/devops/state-of-devops/
You can find me on:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
Music from Uppbeat (free for Creators!).
Intro:
https://uppbeat.io/t/sensho/good-times
License code: QBXDSEGNJZY9DDIC
Outro:
https://uppbeat.io/t/mountaineer/voyager
License code: 5C0VMTUOULFSRSTM
In this episode I share my experience relapsing into anxiety and insomnia, ruminate on an SRE's sphere of influence, and tease an upcoming change of role.
You can find me on:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
Music from Uppbeat (free for Creators!).
Intro:
https://uppbeat.io/t/sensho/good-times
License code: QBXDSEGNJZY9DDIC
Outro:
https://uppbeat.io/t/mountaineer/voyager
License code: 5C0VMTUOULFSRSTM
In this episode I reflect on the book "The Toyota Way" by Jeffrey Liker, and explore four principles which resonate with my work.
The book in question is The Toyota Way: https://www.amazon.com/Toyota-Way-Second-Management-Manufacturer/dp/1260468518
You can find me on:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
Music from Uppbeat (free for Creators!).
Intro:
https://uppbeat.io/t/sensho/good-times
License code: QBXDSEGNJZY9DDIC
Outro:
https://uppbeat.io/t/mountaineer/voyager
License code: 5C0VMTUOULFSRSTM
In this episode I discuss the concept behind continuous delivery and share the ideas we've been exploring at IAG.
The book I mentioned is The Toyota Way: https://www.amazon.com/Toyota-Way-Second-Management-Manufacturer/dp/1260468518
You can find me on:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
Music from Uppbeat (free for Creators!).
Intro:
https://uppbeat.io/t/sensho/good-times
License code: QBXDSEGNJZY9DDIC
Outro:
https://uppbeat.io/t/mountaineer/voyager
License code: 5C0VMTUOULFSRSTM
In this episode I have a chat with Bangser about the transition from testing to SRE, the barriers thrown in front of testers (which SREs don't tend to face), being humble to be let in the door, and much more.
You can find Abby on LinkedIn: https://www.linkedin.com/in/abbybangser/
The book she mentioned was Infrastructure as Code by Kief Morris https://www.thoughtworks.com/insights/books/infrastructure-as-code-2nd-edition
You can find Chastity Majors (cofounder of Honeycomb) on Twitter: https://twitter.com/mipsytipsy
You can find me on:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
Music from Uppbeat (free for Creators!).
Intro:
https://uppbeat.io/t/sensho/good-times
License code: QBXDSEGNJZY9DDIC
Outro:
https://uppbeat.io/t/mountaineer/voyager
License code: 5C0VMTUOULFSRSTM
In this episode I share the story of Grafana Central, an observability platform that we've been standing up at IAG.
You can find me on:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
Music from Uppbeat (free for Creators!).
Intro:
https://uppbeat.io/t/sensho/good-times
License code: QBXDSEGNJZY9DDIC
Outro:
https://uppbeat.io/t/mountaineer/voyager
License code: 5C0VMTUOULFSRSTM
In this episode I share a talk I did earlier in the year as part of the Grafana User Group APAC. I share our experiences attempting to implement SLOs at IAG, and our reliability benchmarking work which is a great way to get started if SRE is brand new to your organisation.
You can find me on:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
Music from Uppbeat (free for Creators!).
Intro:
https://uppbeat.io/t/sensho/good-times
License code: QBXDSEGNJZY9DDIC
Outro:
https://uppbeat.io/t/mountaineer/voyager
License code: 5C0VMTUOULFSRSTM
In this episode I share experiences and ideas about Kubernetes, and what I learned from speaking to Ruben Hakopiean from Kubevious.
I'd like to give a huge shout out to Ruben. Many of the topics and ideas discussed come straight from what was discussed in the interview we recorded (but were unable to publish due to audio issues).
You can find Ruben on LinkedIn: https://www.linkedin.com/in/rubenhak/
And find out more about Kubevious here: https://kubevious.io/
You can find me on:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
Music from Uppbeat (free for Creators!).
Intro:
https://uppbeat.io/t/sensho/good-times
License code: QBXDSEGNJZY9DDIC
Outro:
https://uppbeat.io/t/mountaineer/voyager
License code: 5C0VMTUOULFSRSTM
In this episode I have a chat with Joey Hendricks about running performance tests in production.
You can find Joey on LinkedIn: https://www.linkedin.com/in/joey-hendricks/
And GitHub: https://github.com/JoeyHendricks
You can find me on:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
Music from Uppbeat (free for Creators!).
Intro:
https://uppbeat.io/t/sensho/good-times
License code: QBXDSEGNJZY9DDIC
Outro:
https://uppbeat.io/t/mountaineer/voyager
License code: 5C0VMTUOULFSRSTM
In this episode I share my takeaways from the NZ DevOps Summit held in Auckland. This was the first in-person event I had attended in three years.
You can find me on:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
Music from Uppbeat (free for Creators!).
Intro:
https://uppbeat.io/t/sensho/good-times
License code: QBXDSEGNJZY9DDIC
Outro:
https://uppbeat.io/t/mountaineer/voyager
License code: 5C0VMTUOULFSRSTM
In this episode I have a chat with Chris Evans from incident.io about using incidents to lift the lid on an organisation, how aiming for zero incidents can stall an organisation, how tracking MTTR is unhelpful, and much more.
You can find Chris on LinkedIn: https://www.linkedin.com/in/evnsio/
Here are the resources Chris mentioned...
The practical guide to incident management:
http://incident.io/guide
The Field Guide to Understanding Human Error (by Sidney Dekker) https://www.oreilly.com/library/view/the-field-guide/9781317031833/
Moving Past Shallow Incident Data by John Allspaw https://www.adaptivecapacitylabs.com/blog/2018/03/23/moving-past-shallow-incident-data/
And the SREcon talk he mentioned by Courtney Nash: https://www.usenix.org/conference/srecon22americas/presentation/nash
You can find me on:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
Music from Uppbeat (free for Creators!).
Intro:
https://uppbeat.io/t/sensho/good-times
License code: QBXDSEGNJZY9DDIC
Outro:
https://uppbeat.io/t/mountaineer/voyager
License code: 5C0VMTUOULFSRSTM
In this episode I have a chat with Ganesh Datta, CTO and co-founder of Cortex.io. In this episode we discuss the human challenges of microservices, gamifying reliability, connecting business outcomes with SRE work, and much more.
You can find Ganesh on LinkedIn: https://www.linkedin.com/in/gsdatta/
You can find me on:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
Music from Uppbeat (free for Creators!).
Intro:
https://uppbeat.io/t/sensho/good-times
License code: QBXDSEGNJZY9DDIC
Outro:
https://uppbeat.io/t/mountaineer/voyager
License code: 5C0VMTUOULFSRSTM
In this episode I have a chat with Sebastian Vietz, an SRE lead based in Canada who has been leading the implementation of SRE across different teams and organisations for eight years. In this episode we discuss SLO adoption, SRE going mainstream, virtual teams, and many other topics.
You can find Sebastian on LinkedIn: https://www.linkedin.com/in/sebastianvietz/
You can find me on:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
Music from Uppbeat (free for Creators!).
Intro:
https://uppbeat.io/t/sensho/good-times
License code: QBXDSEGNJZY9DDIC
Outro:
https://uppbeat.io/t/mountaineer/voyager
License code: 5C0VMTUOULFSRSTM
In this episode I discuss potential pre-requisites that are ideally in place before attempting to adopt SLOs.
You can find me on:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
Music from Uppbeat (free for Creators!).
Intro:
https://uppbeat.io/t/sensho/good-times
License code: QBXDSEGNJZY9DDIC
Outro:
https://uppbeat.io/t/mountaineer/voyager
License code: 5C0VMTUOULFSRSTM
In this episode I share my updated thinking on SLOs and an "ah-ha!" moment I had.
You can find me on:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
Music from Uppbeat (free for Creators!).
Intro:
https://uppbeat.io/t/sensho/good-times
License code: QBXDSEGNJZY9DDIC
Outro:
https://uppbeat.io/t/mountaineer/voyager
License code: 5C0VMTUOULFSRSTM
What is latency and how does it relate to customer experience? Where do you measure it? Why do the metrics we choose to capture matter?
You can find me on:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
Music from Uppbeat (free for Creators!).
Intro:
https://uppbeat.io/t/sensho/good-times
License code: QBXDSEGNJZY9DDIC
Outro:
https://uppbeat.io/t/mountaineer/voyager
License code: 5C0VMTUOULFSRSTM
When it comes to reliability SLO's and NFR's are both somewhat related in that they allow us to describe the level of service we want to provide our customers. So how do they match up head to head?
You can find me on:
LinkedIn: https://www.linkedin.com/in/stephento...
Twitter: https://twitter.com/the_kiwi_sre
Music from Uppbeat (free for Creators!).
Intro:
https://uppbeat.io/t/sensho/good-times
License code: QBXDSEGNJZY9DDIC
Outro:
https://uppbeat.io/t/mountaineer/voyager
License code: 5C0VMTUOULFSRSTM
What is an error? Where do you measure errors? How do they relate to SLO's and error budgets?
You can find me on:
LinkedIn: https://www.linkedin.com/in/stephento...
Twitter: https://twitter.com/the_kiwi_sre
Music from Uppbeat (free for Creators!).
Intro:
https://uppbeat.io/t/sensho/good-times
License code: QBXDSEGNJZY9DDIC
Outro:
https://uppbeat.io/t/mountaineer/voyager
License code: 5C0VMTUOULFSRSTM
In this episode we discuss the observability concept of a 'single pane of glass' view, and I share my experience implementing one.
You can find me on:
LinkedIn: https://www.linkedin.com/in/stephento...
Twitter: https://twitter.com/the_kiwi_sre
Music from Uppbeat (free for Creators!).
Intro:
https://uppbeat.io/t/sensho/good-times
License code: QBXDSEGNJZY9DDIC
Outro:
https://uppbeat.io/t/mountaineer/voyager
License code: 5C0VMTUOULFSRSTM
In this episode I share my thoughts from SLOconf, a conference all about Service Level Objectives (SLO's).
You can find all the talks (actually 60 of them!) from SLOconf here: https://www.youtube.com/watch?v=pgZm2Bp2-AQ&list=PLLNq9CBV7AFwkXvYmjPPIQlRDVwTmacEK
You can find me on:
LinkedIn: https://www.linkedin.com/in/stephento...
Twitter: https://twitter.com/the_kiwi_sre
Music from Uppbeat (free for Creators!).
Intro:
https://uppbeat.io/t/sensho/good-times
License code: QBXDSEGNJZY9DDIC
Outro:
https://uppbeat.io/t/mountaineer/voyager
License code: 5C0VMTUOULFSRSTM
In this episode I provide three more (p)re-reactions to upcoming sessions at o11yfest 2022 (a conference all about observability). The talks I cover are:
"Obserability driven development" by Jessica Kerr
"Return on investment driven observability" by Michael Hausenblas
"How the OpenTelemetry Collector puts you in the driver seat" by Alex Boten
I am also speaking at o11yfest. You can watch my talk on Bad Observability at o11yfest from May 9th to the 12th: https://o11yfest.org/
You can find me on:
LinkedIn: https://www.linkedin.com/in/stephento...
Twitter: https://twitter.com/the_kiwi_sre
Music from Uppbeat (free for Creators!).
Intro:
https://uppbeat.io/t/sensho/good-times
License code: QBXDSEGNJZY9DDIC
Outro:
https://uppbeat.io/t/mountaineer/voyager
License code: 5C0VMTUOULFSRSTM
In this episode I provide a (p)re-reaction to three of the talks that will be included in the upcoming 2022 o11yfest conference (a conference all about observability). The talks I cover are:
"Where the heck are my spans?" by Reese Lee
"Confidence in chaos" by Narmatha Bala
"Is MTTR still relevant in a modern, cloud native world?" by Martin Mao
I am also speaking at o11yfest. You can watch my talk on Bad Observability at o11yfest from May 9th to the 12th: https://o11yfest.org/
You can find me on:
LinkedIn: https://www.linkedin.com/in/stephento...
Twitter: https://twitter.com/the_kiwi_sre
Music from Uppbeat (free for Creators!).
Intro:
https://uppbeat.io/t/sensho/good-times
License code: QBXDSEGNJZY9DDIC
Outro:
https://uppbeat.io/t/mountaineer/voyager
License code: 5C0VMTUOULFSRSTM
How do you measure the availability of your services? What metric do you pick? What layer of the solution do you track it from?
My colleague Gwen and I are sharing our SLO definition workshop experience at SLOconf from May 9th to 12th: https://www.sloconf.com/
You can also watch my talk on Bad Observability at o11yfest from May 9th to the 12th: https://o11yfest.org/
You can find me on:
LinkedIn: https://www.linkedin.com/in/stephento...
Twitter: https://twitter.com/the_kiwi_sre
Music from Uppbeat (free for Creators!).
Intro:
https://uppbeat.io/t/sensho/good-times
License code: QBXDSEGNJZY9DDIC
Outro:
https://uppbeat.io/t/mountaineer/voyager
License code: 5C0VMTUOULFSRSTM
In the episode I share our team's experience defining SLO's, and how we experimented and pivoted to achieve better outcomes.
As discussed in the episode, if you would like to hear (and see) more about our SLO workshop, my colleague Gwen and I are speaking at SLOconf from May 9th to 12th: https://www.sloconf.com/
You can find me on:
LinkedIn: https://www.linkedin.com/in/stephento...
Twitter: https://twitter.com/the_kiwi_sre
Music from Uppbeat (free for Creators!).
Intro:
https://uppbeat.io/t/sensho/good-times
License code: QBXDSEGNJZY9DDIC
Outro:
https://uppbeat.io/t/mountaineer/voyager
License code: 5C0VMTUOULFSRSTM
What are even more antipatterns to avoid in monitoring, alerting, tracing, and logging?
Shout out to James Pulley for his contribution to this episode. James is one of the world's leading experts on performance engineering and can be found on LinkedIn here: https://www.linkedin.com/in/jameslpulley3/
I will be presenting about Bad Observability at o11yfest from May 9th to the 12th 2022: https://o11yfest.org/
You can find me on:
LinkedIn: https://www.linkedin.com/in/stephento...
Twitter: https://twitter.com/the_kiwi_sre
Music from Uppbeat (free for Creators!).
Intro:
https://uppbeat.io/t/sensho/good-times
License code: QBXDSEGNJZY9DDIC
Outro:
https://uppbeat.io/t/mountaineer/voyager
License code: 5C0VMTUOULFSRSTM
What are some more antipatterns to avoid in monitoring, alerting, tracing, and logging?
Shout out to Raguraman Balasubramanian (https://www.linkedin.com/in/raguraman-balasubramanian-070150108/) for his contribution to this episode.
You can find me on:
LinkedIn: https://www.linkedin.com/in/stephento...
Twitter: https://twitter.com/the_kiwi_sre
Music from Uppbeat (free for Creators!).
Intro:
https://uppbeat.io/t/sensho/good-times
License code: QBXDSEGNJZY9DDIC
Outro:
https://uppbeat.io/t/mountaineer/voyager
License code: 5C0VMTUOULFSRSTM
What are some antipatterns to avoid in monitoring, alerting, tracing, and logging?
Music from Uppbeat (free for Creators!).ย
Intro:
https://uppbeat.io/t/sensho/good-times
License code: QBXDSEGNJZY9DDIC
Outro:
https://uppbeat.io/t/mountaineer/voyager
License code: 5C0VMTUOULFSRSTM
What is SRE really about? How did it start? What do I want it to be? What is it being implemented as in the industry?
Music from Uppbeat (free for Creators!).ย
Intro:
https://uppbeat.io/t/sensho/good-times
License code: QBXDSEGNJZY9DDIC
Outro:
https://uppbeat.io/t/mountaineer/voyager
License code: 5C0VMTUOULFSRSTM
In this episode I wrap up the Performance Time show with some commentary around my blog "Wrapping up 13 years as a Performance Engineer" (https://www.linkedin.com/pulse/wrapping-up-13-years-performance-engineering-stephen-townshend/)
This week we look at whether to identify SLI's or SLO's first - and take a look at how both approaches might look like with a ridiculous example.
The Google Site Reliability Book is free here: https://sre.google/books/ (I bought and listened to the audiobook on Audible).
The intro and background music is "Elevator Music Lofi" by Oleksii Kaplunskyi.
What are SLI's, SLO's, and SLA's? Why do they matter? How can you identify SLI's for your products?
The intro and background music is "Elevator Music Lofi" by Oleksii Kaplunskyi.
How do you know what to alert on? And are there things we shouldn't alert on? What about dashboards?
In this episode we talk about locating signals amongst the noise of monitoring and logging data in our organisations.
This week we talk about the Four Golden Signals, a starting point when choosing what to monitor and alert on in your software platforms.
In this episode I share my first week working as an SRE and try to explain what it's all about.
In this update I give a quick update on how things are going and where the podcast is going next.
In this episode I share an experience of work related anxiety I recently went through.
Shout out to Joey Hendricks for reviewing this content before it went live.
What are the challenges when performance engineers are asked to move away from hands on technical work, into advocating and enabling others?
Shout out to Sajeesh Nair and Ben Rowan who's LinkedIn comments I refer to.
In this episode we talk about measuring whether performance itself has a meaningful impact on customer experience, and how difficult it can be to answer that question.
Third party trackers
Are crowding my website
Analytics requests
Are bringing me down
Guess Iโll just close my tab
Oh yeah
Alright
Feels bad
Inside
Four hundred requests
For just one web page
My mobile netwo-oo-er-oo-erk
Is struggling to cope
Iโll guess Iโll call instead
Say it ainโt slow
Your site is unbearable
Say it ainโt slow
My business is going elsewhere
I canโt connect to
Anything you do
All of your websites
Are timing out
When I try
To pay for the product that you had up on
Your website just the other day
Itโs not cool
Say it ainโt slow
Your site is unbearable
Say it ainโt slow
My business is going elsewhere
Dear website
I try you, in spite of years of slowness
Youโve cleaned up, got faster, (even) using a CDN now
This gigan-tic image, awakens ancient feelings
Frustration, desperation, all of your customers are goooooonneeeโฆ
Yeah, yeah, yeah, yeah, yeah
Say it ainโt slow
Your site is unbearable
Say it ainโt slow
My business is going elsewhere
Doing more early (shifting left) and doing more in production (shifting right) to reduce how much end to end performance testing we need to do.
In this episode I chat with globally recognized performance specialist Nicole van der Hoeven about chaos engineering, diversity in technology, note taking, chasing shiny things, having a job interview with Stijn Schepers, and much more.
Nicole can be found on Twitter @n_vanderhoeven or e-mail nicole@k6.ioย
The two books mentioned are Modern Technical Writing by Andrew Etter and The Art of Statistics by David Spiegelhalter.
In this episode we look at performance testing in stubbed environments versus fully integrated end to end environments. We drill into some of the benefits and limitations of stubbed component performance testing, and how to make sense of it all in our organisations.
In this episode we explore the unique challenges of managing test data during load testing, and some solutions to make our lives easier.
Join Stijn and I as we discuss culture change, coincidences, war rooms, unicorns, penguins, McDonald's testing, and many other things.
Join me as a I chat with performance specialist Alan Gordon about stress and anxiety as a performance engineer.
How do we present ideas about performance engineering either in meetings or formal presentations? And I tell a story about having an anxiety attack in a meeting.
How to do the great technical work you do justice by making sure the right people hear about it in a way they understand.
Join me as a I chat with performance specialist Ben Rowan about his collaborative and scientific style of performance engineering, the challenges of cloud performance, and the time he wiped five years of work.
A look at the different styles or kinds of performance engineer. Which kinds are you? What areas are you interested in growing?
How can we grow into better performance engineers? How do we avoid getting stuck in our ways?
A look at the some of the ways we can deal with stress and anxiety as a performance engineer.
A look at the contributing factors of why performance testing and engineering can be so stressful.
A real world example of monitoring the performance of a containerized platform. Includes some example behaviours to watch out for.