Welcome to Ship Happens where we interview interesting people in the DevOps and engineering community and get real about their careers, paths to DevOps and all the highs and lows in between.
How to effectively communicate asynchronously
In episode 7 of Ship Happens, Sam Goldstein explains what it means to communicate with a global team in an asynchronous fashion. While 9-to-5 offices typically expect real-time responses to Slack messages, Sam explains what it means when you're constantly working with people on the other side of the world. See how GitLab prioritizes their meetings and consistently documents human workflows in order to help teams work asynchronously with one another.
Maintaining the human element while working remotely
See how and why Sam and the GitLab team put an emphasis on building stronger personal connections over video calls. Understand how remote teams can bolster morale and tighten relationships between teammates to drive more productive teams.
When to use Zoom, Slack, Emojis, etc.
GitLab's remote handbook covers all types of communication workflows. What type of channel should be used for which types of work? GitLab is the world's largest fully-distributed team, allowing Sam to offer great tips and tricks for when to communicate in certain ways and what to expect when communicating with a global engineering team.
What do you do when S3 goes down?
Hear about Sam's on-call horror story, when nearly half the internet went out. Learn more about how teams can approach "100-year flood" scenarios and how to balance reliability with speed, improving customer experiences as well as on-call workflows.
Reading material mentioned in the show
GitLab's remote work resources: https://about.gitlab.com/company/culture/all-remote/resources/
GitLab's remote handbook: https://about.gitlab.com/handbook
Maintaining availability and velocity in a highly-regulated environment
Tammy talks about what it takes to maintain available, secure services in a highly-regulated environment. See how teams think about their delivery pipelines and services when applications and infrastructure need to adhere to strict Australian governmental regulations.
Tammy's path to SRE and the organizational value of SRE
Through personal experiences, Tammy discovers the value of SRE in a very real way. Tammy talks about why this piqued her interest in site reliability engineering and how she made the move from a full-stack engineer to an SRE. She then elaborates on her journey into SRE and talks about how managers and engineers can get organizational buy-in for SRE and show the value of it over time.
How to train on-call teams for incident response
Incident management, real-time response and on-call efficiency are important to Tammy, and should be for SREs everywhere. Tammy will cover actionable tips for training on-call teams and giving on-call responders the tools and resources they need to make on-call suck less. Tammy also dives into her expertise in skateboarding and how some of the things she learned while skateboarding has made her and her teams better.
Chaos engineering and real applications of it
Tammy discusses the topic of chaos engineering and the intentional injection of failure into your systems – so you can learn from it and make your systems more resilient over time. Through tabletop exercises, gamedays, on-call training and post-incident reviews, Tammy shows how teams can improve both people operations and technical operations around incident response
Reading material mentioned in the show
Tammy's O'Reilly Book: Reducing MTTA for High-Severity Incidents: https://www.oreilly.com/library/view/reducing-mttd-for/9781492046202/Gremlin's Chaos Engineering Slack Community: https://www.gremlin.com/slack/
Gremlin's Resources for Site Reliability Engineering: https://www.gremlin.com/site-reliability-engineering/
Tammy's Twitter Account: https://twitter.com/tammybutow
Hear Andi's story and his contributions to the State of DevOps Report
Andi Mann has worked in nearly all departments of IT, software engineering and advocacy in his 35+ years in the industry. Now, Andi leads the innovation research and advisory team for Splunk’s IT Markets business, helping support leading-edge customers with DevOps, SRE, AppSec, CI/CD etc., as well as promoting the numerous use cases for Splunk itself. He uses his expertise and his background to learn about how the DevOps industry is evolving and publishes his findings in the 2019 State of DevOps Report.
How to do this DevOps 'thing'
Andi gives some great advice for getting organization buy-in for DevOps and SecOps practices and how to tightly integrate them without hindering velocity. Implementing DevOps isn't as simple as buying a tool and automating a process or two. Andi describes some real ways to think about DevOps in order to help any team, at any level of maturity, adopt DevOps in a thoughtful way that drives value.
Why Ops is the center of your business
With developers managing their own code and taking accountability for production uptime and on-call responsibilities, developers are now taking part in operations functions (i.e. DevOps). See how Ops contributes value to all aspects of the business, including sales, engineering, product, etc. If your applications and infrastructure experience downtime or frequent performance errors, they negatively impact customer sentiment and experience - hurting your business.
On-call during a London hurricane
As an on-call contractor, living near the affected data center, what happens when a hurricane sweeps through London? With no help coming, with just a few people managing a major incident in the midst of hurricane recovery, Andi shows just how bad on-call can suck.
Gene's New Book: The Unicorn Project
Gene Kim discusses his re-exploration of the Parts Unlimited universe from a different angle and how he came up with the Five Ideals that he surfaces in the book. Benton and Gene take a deep dive into the problems the book's protagonist, Maxine, faces in her DevOps journey, the reasons why those problems exist, and actionable solutions.
Encouraging Psychological Safety and an Inclusive DevOps Culture
Explore psychological safety, leadership and building positive, DevOps-minded team chemistry through fun anecdotes about the San Antonio Spurs and General Stanley McChrystal. Find out how any team can drive efficient engineering workflows through empowerment, inclusivity and collaboration.
Using Procedural vs. One-Shot Learning in Engineering to Address Business Concerns
Why is it important for engineering teams to focus on work that solves business problems? How do you build a process for continuous improvement that helps encourage functional programming and ultimately delivers reliable, secure applications and services faster?
On-Call Horror Story: How a Bad Perl Script Led to Starting a Company
Hear about Gene's experience working in a computing center in 1991 and how one bad Perl script eventually led him into starting Tripwire.
From Developer to Growth Engineering
Kim talks about her 20+ year working in software – from developer to growth engineering, with a sprinkle of UX and product management too. Learn how Kim’s broad skill set drives powerful data-driven decisions and helps customers realize product value faster.
The Rotating Flywheel of Software Development and Product Marketing
Kim and Benton discuss the concept of a rotating flywheel for software engineering, testing, deployment and product marketing. They then discuss how important it is to keep the cycle moving in order to not only continuously deliver code to production, but to quickly show this initial product value to customers.
Building a “Welcome Experience” in SaaS Applications
See how JumpCloud has approached a product-forward approach to growth engineering and marketing with their “welcome experience”. Kim talks about the value of customer-focused collaboration between engineering and business teams when building SaaS applications.
On-Call Horror Story: Brute Force Attack on a Saturday?
Did a change to the verification process cause an outage? Nope. There was a brute force attack on a Saturday, of course. See how multiple teams came together to quickly respond to the incident and turn an on-call horror story into an event of the past.
The (Frustrating) Path to DevOps
Like many developers, Patrick’s path to DevOps was paved with frustration in slow, inefficient processes. Now, Patrick specializes in teaching DevOps to team members.
What it Means to Be Truly Bootstrapped
Raygun doesn’t just throw money at problems. The company takes a strategic approach to solving all issues, whether development or team lunch related. In order to iterate and deliver customer value quickly, Raygun is their own customer.
On-Call Horror Story: The Twitter Emoji No One Saw Coming
Learn how a firework emoji crashed Patrick’s application, and how he got through the firefight.
Transforming On-Call at NS1
Like many companies, NS1 struggled to create a humane on-call culture for the team responsible for keeping its services up and running 24/7. When Bethany Abbott was hired to lead the TechOps team, she set out to transform the on-call culture by making strategic hires, tweaking on-call rotations and enlisting the right technology.
Getting Singled Out as a Woman in Tech
Long before Bethany was running TechOps for major tech companies, she experienced inequality in the tech industry first hand, as the only woman in all of her computer science classes. Now, she’s on a mission to introduce future generations of women and minorities to technology early on through her organization, Future Tech Women.
On-Call Horror Story: Putting a Country Back Online…Over Christmas
Not many people can say they put an entire country back online, but Bethany did, at the expense of her Christmas holiday. Find out what this experience was like, and how she restructured her team’s on-call protocol to ensure it didn’t happen again.