Cloud Native Startup: Recent Episodes

Emily Omier

Are you an entreprenurial software engineer who wants to learn more about how to go from being a highly accomplished software engineer to starting and growing a company that makes products aimed at engineers in the cloud native ecosystem? This podcast has stories from founders who have made the transition from engineer to startup founder as well as advice from the tech professionals that help founders succeed.

View Details

Philippe Humeau is the CEO and Co-Founder of CrowdSec, an open-source security company with a very unique business model that doesn’t fit the usual open source patterns. Philippe talked about how to focus on providing a fair exchange of value between maintainers / open source companies and users, and how to monetize a project that is providing value for free.

Philippe also talked about why he thinks open-source founders are under more pressure to get their business model right at the start, tips on making the right hiring decisions, and how to communicate with the community in an effective and transparent way. I also liked Philippe’s cynicism: why he views open source as primarily a pragmatic choice for his business, given the type of company he wanted to build.

Philippe also shares the logic behind his uncommon view that only making certain features available to paying customers isn’t a truly open-source business strategy.

Highlights:

  • I introduce Philippe, who gives some background on his career journey and what he does at CrowdSec (00:22)
  • Philippe explains why it seems that security companies are underrepresented in the open-source space (03:19)
  • The most common mistake Philippe sees when people start an open-source business (05:03)
  • Why Philippe believes that open-source companies are under more pressure to get their business model right the first time (09:26)
  • How Philippe came up with Crowdsec’s unique business model (16:15)
  • The pushback that Philippe got when he presented his business model initially (19:33)
  • Why Philippe views open source as a means to an end, and how that has affected his choices at CrowdSec (25:10)
  • The most interesting mistake Philippe has made since starting CrowdSec (27:28)
  • Why Philippe believes open source business models are more promising than closed source (31:19)
  • The advice that Philippe would give to an open source founder who is looking to build a successful company (34:11)
  • Why Philippe feels that having certain features behind a paywall is not a truly open-source business model (35:53)
  • Where you can learn more about Philippe and connect with CrowdSec (40:11)

Links:

Philippe

  • LinkedIn: https://www.linkedin.com/in/philippehumeau/
  • Twitter: https://twitter.com/philippe_humeau
  • Company: https://www.crowdsec.net/

View Details

Kevin Muller is the CEO and co-founder of Passbolt, a security-first, open-source password manager, and he joined me to talk about the risks of having too much time and money, the value of getting trashed on social media and why he values in-person interactions with the team.

There were a lot of interesting pieces to pick apart from this episode. First of all, Kevin talked about the importance of not commercializing too early. I think he's the only founder I've ever heard say something along those lines, but he makes a good argument. (Also, Tim Chen and I talked about the timing of commercialization last week, my takeaway is that no one feels like they commercialized at precisely the right moment). Second, we had a good discussion about the different priorities of European versus American investors can push companies to make different decisions. The subtext that we didn't address directly is make sure you are aware that your investors priorities are going to influence how your company evolves, choose your investors with that in mind. (and check out the episode with Markus Düttmann if you want more on the EU vs US investment environment for open source startups). Lastly, password managers have been in the news, and not in a good way — and how to best react to a super embarrassing situation for a competitor is not always obvious. So we talked about how Passbolt has tried to steer the conversation about password management in light of recent high-profile hacks in the ecosystem.

Highlights:

  • Kevin introduces himself and describes his work at Passbolt (00:26)
  • How Kevin got the idea for Passbolt and the story of how he brought his idea to life (01:07)
  • The mistakes that Kevin and his co-founders made when launching Passbolt (05:03)
  • What happened when Kevin and his co-founders officially launched Passbolt in 2016 (08:12)
  • How Kevin and his co-founders decided to move from a purely open-source product to a commercialized product (09:32)
  • Why Passbolt is a hybrid company and the value Kevin sees in having employees spend time in the office (12:41)
  • Kevin describes why it was so important for Passbolt to be an open-source company (15:58)
  • Why Kevin feels it’s important not to commercialize an open-source product too quickly (19:07)
  • The different priorities of European VCs versus U.S. VCs (21:58)
  • Why honest feedback is so valuable and how Kevin and his team evaluated the feedback they got at the launch of Passbolt (24:26)
  • Kevin’s reaction to data breaches that happen to other password management solutions (27:09)
  • The biggest challenges that Kevin and the team at Passbolt are working on currently (31:09)
  • Kevin’s advice to open-source founders (32:47)

Links:

  • LinkedIn: https://www.linkedin.com/in/kevinmuller80/
  • LinkedIn: https://www.linkedin.com/company/passbolt/
  • Twitter: https://twitter.com/passbolt
  • Company: https://www.passbolt.com

View Details

Tim Chen is a Partner at Essence VC and also the Co-Host of the Open Source Startup Podcast. Through these channels, he has the opportunity to speak with a broad variety of open source startups. Throughout our conversation, we explore the patterns that Tim sees in the open source startup space. Tim talked about how too many founders take the decision to build an open source company too lightly and the path that he would take if he were to start a venture-backed open source startup tomorrow. We also discuss the different monetization models of open-source startups and the true business value of an open source project.

Highlights:

  • Tim introduces himself and describes his role at Essence VC as well as his work as Co-Host of the Open Source Startup Podcast (00:22)
  • The common patterns that Tim sees having worked with so many open source startups (02:25)
  • Tim describes the landscape of open source and how it varies from open source projects to venture-backed, open source companies (06:48)
  • What path Tim would take if he were to start a venture-backed, open source startup tomorrow (09:31)
  • How Tim views different monetization models and their potential profitability (17:29)
  • Tim’s views on the pros and cons of an open-core model (20:34)
  • The business value of an open source project according to Tim (24:47)
  • How Tim’s evaluation and investing tactics have changed as he’s worked with more open source startups (31:58)
  • Where listeners can find more information about Tim and learn more about his work (37:47)

Links:

Tim

  • LinkedIn: https://www.linkedin.com/in/timchen
  • Twitter: https://twitter.com/tnachen
  • Company: https://www.essencevc.fund/

View Details

Franz Karlsberger is the CEO of Amazee.io, an open-source platform that seeks to make developers’ lives easier by abstracting their day-to-day workload. Throughout our conversation, we explore what it means to join an open-source start-up as an external CEO, and why Franz put so much emphasis on go-to-market strategy. Franz also walks through the importance of knowing what open-source business model your company will follow, and how to measure the success of an open-source project.

Listen in as Franz shares some of his most interesting mistakes, what he’d do differently if he could start over, and why Franz feels open-source is more than just a type of software, it’s a company ethos that affects everything down to the team culture.

Highlights:

  • Franz introduces himself and his company Amazee.io, which is a ZeroOps application delivery platform (00:50)
  • How Amazee.io went from being a point solution to a platform solution (06:20)
  • Why Franz was brought in as an external CEO for Amazee.io to accelerate growth (10:03)
  • How Franz adjusted to working at an open-source start-up and what that learning curve was like for him (11:47)
  • The importance of open-source at Amazee.io and why it is baked into their core values and ethos as a company (15:30)
  • How the go-to-market model differs for Amazee.io’s cloud offering versus their managed offering (17:51)
  • Franz describes some of the most interesting mistakes he’s made and what he’s learned from them (23:25)
  • Franz’s views on measuring the success of an open-source project (26:29)
  • How listeners can connect with Franz and learn more about Amazee.io (32:37)

Links:

Franz

  • LinkedIn: https://www.linkedin.com/in/franzkarlsberger/
  • Twitter: https://twitter.com/fkarlsberger
  • Company: https://www.amazee.io/

View Details

It’s kind of a cliche, Vlad A. Ionescu, founder and CEO of Earthly, says, but his first attempts to build something really awesome focused on amazing technology. With hindsight, he doesn’t think it’s so surprising that those efforts weren’t successful. It’s not that passion doesn’t matter, but rather that he had to learn to build things that inspired passion from both the market and the builders. We also talked about:

  • Leaving a job, blowing through his savings, going back to a job before finding entrepreneurial success
  • Realizing that if he wanted to have the kind of impact on the world that he wanted to, he had to figure out a way to make it as an entrepreneur, because the alternative was climbing the corporate ladder and that didn’t sound like fun
  • Why it’s important to be brutally honest with yourself and what you suck at
  • How Vlad finally found success at Shift Left (now Quiet.ai)

I also really liked his ideas about cutting corners — that startups will always have to cut some corners, it’s just up to you to decide which ones to cut.

Highlights:

  • Vlad recounts lessons learned from early entrepreneurial failures. (2:31)
  • Taking failure personally to overcome weaknesses (5:19)
  • Vlad explains what led to his first success with Shift Left (7:40)
  • Vlad shares his journey from Shift Left to Earthly (13:40)
  • Why open source? (17:03)
  • How Vlad and his team built Earthly based on what he learned from building Shift Left (25:07)
  • Breaking a product down to its components to find more value (31:38)
  • The Startup Hierarchy of Needs (34:02)

Links:

Vlad

  • LinkedIn
  • Twitter
  • Earthly
  • Qwiet.ai
  • Personal Site

View Details

Markus Düttman, a former Principal at Nauta Capital, is steeped in the European open source scene. From his beginnings in theoretical physics, Düttman’s hard pivot into venture capital funding granted him a spot in the developing tech world as a connoisseur of the culture and a champion for start-ups. He even contributed to Nauta Capital’s European Open Source Report detailing the state of the ecosystem as of October 2022.

On this episode of the Business of Open Source, listen to his insight into the European markets, the various business models generally used for open source start-ups, and what he looks for in an open source start-up.

Note: Markus has since left Nauta and is on paternity leave. He also asked me to add a follow-up to the episode: After thinking more about the biggest danger to open source companies, he thinks most of them will fail from problems like not building the right team, failure to find product market fit and/or failure to monetize. Hyperscalers are a danger, but probably won’t be what causes most startups to fail.

Highlights:

  • The unique qualities of the European open source ecosystem (2:20)
  • European market vs. the American Market in terms of funding (3:45)
  • Advantages and disadvantages of a European open source company (4:40)
  • Navigating the use of different business models within a business (8:14)
  • How Markus evaluates open source start-ups (11:46)
  • Don't be the open source version of an existing enterprise company (16:14)
  • Signs of a company worth investing in (16:45)
  • Potential risks to the open source ecosystem in the coming years (19:24)

Links:

Markus

  • LinkedIn
  • Twitter
  • Nauta Capital
  • European Open Source Report

View Details

This week Shanea Leven, CEO and Co-Founder of CodeSee, joins me to chat about demystifying code bases and building an effective team.

In this episode of The Business of Open Source, Shanea and I discuss the origins of her company, CodeSee, how it morphed from a training course to a SaaS product, and how they contribute to open source even though CodeSee is not an open source company. Shanea also shares valuable insight into working closely with your spouse, the importance of communication and empathy in building an effective team, and how she’s evolved as a leader. Listen to hear all of her insight and advice, and find out how the CodeSee SaaS offering helps companies understand their code bases and make critical decisions faster.

Highlights:

  • How Shanea’s experience at Docker influenced her decision about whether or not to make CodeSee open source (1:51)
  • Origin and history of CodeSee (6:58)
  • How CodeSee progressed from a training course to a SaaS product (10:41)
  • Shanea’s advice for other entrepreneurs interested in founding a startup with their spouse (14:36)
  • Lessons Shanea has learned from the challenges of building a team (17:18)
  • The importance of face time with customers (23:08)
  • How CodeSee works (26:31)
  • How CodeSee contributes to open source without being open source (30:14)

Links:

Shanea

  • LinkedIn
  • Twitter
  • CodeSee

View Details

This week Heikki Nousiainen, CTO and Co-founder of Aiven, joins me to chat about building the business, his passion for open source and entrepreneurship, and his hopes for the future of open source in the public sector.

In this episode, Heikki and I explore the successes and challenges he and his three co-founders encountered in creating and maintaining their global open source data platform. We discuss how they choose technologies to support, the importance of customer demand, how founders can learn to work together, and when to “kill your darlings.”

Highlights:

  • Origins of Aiven (1:40)
  • Pros and cons of being headquartered in Helsinki (4:41)
  • Aiven’s relationship to the open source community (6:02)
  • How Aiven has evolved since its inception (7:34)
  • How Aiven chooses technologies to incorporate into their service offerings (9:21)
  • One thing that has been very successful for Aiven (12:51)
  • Why Aiven chose their business model (16:33)
  • The biggest challenge Aiven is currently facing (17:37)
  • The State of Open Con and giving back to the Community (20:17)
  • Barriers to more open source adoption in the public sector (21:24)

Links:

Heikki

  • LinkedIn
  • Twitter: @hnousiainen
  • Aiven

View Details

Michael Cheng, Chief Legal Officer at Aalyria Technologies, is a master at strategy and execution for open-source products and companies. From his humble beginning spearheading the open source team at Meta (formerly Facebook), Cheng has honed his knowledge about the interworking of open source and utilizes it to its fullest potential.

In this episode of The Business of Open Source, Cheng talks about his time as Meta's lead in open source as well as what it's like to be an individual working for a large company. He also explains what happens in mergers and acquisitions with open source projects and the legality of being a small fish in a large pond!

Highlights:

  • A little insight into Meta's open source (2:06)
  • Detractors on a project (4:25)
  • It's hard running a large company open source (7:47)
  • Is it a problem to be individually driven? (9:45)
  • Creating projects while working for large companies in open source (13:25)
  • Do they have a right to reprimand? (16:16)
  • What happens in a merger for an open source company (17:53]
  • Recommendations to an inquirer (21:47)
  • Personality-deprived communities (24:48)

Links:

Michael

  • LinkedIn
  • Aalyria

View Details

Are you struggling to find a co-founder? Having trouble navigating a relationship with your partners? These are all questions Tanis Jorge, CEO of The Co-founder's Hub, tackles daily in her work. A serial tech entrepreneur and a leading entrepreneur advisor, it is no wonder Jorge has founded and built many businesses, such as Trulioo and IQuiri Inc.

On this episode of The Business Of Open Source, I ask Jorge about finding the right co-founder, why legal frameworks are so important, and the hi's and lows of collaboration. We also discuss warning signs, life stages, and why everyone is a janitor!

Highlights:

  • The right co-founder (2:25)
  • Life stages are important (4:31)
  • The uncomfortable task (5:50)
  • The legal framework (8:00)
  • Everybody's the janitor (10:05)
  • The hi's and lows of collaboration (14:46)
  • Navigating problems and relationships(18:45)
  • No one is talking about this (23:24)
  • Meeting a bad partner (24:35)
  • Multiple founder relationships (26:10)

Links:

Tanis

  • LinkedIn
  • Twitter
  • The Cofounder's Hub
  • Trulioo

View Details

From college dropout to developer efficiency guru, Kyle Campbell knows his way around workflow integrations for software companies. He is a fierce proponent of efficient workflows and hopes to spread his experience to any company that can benefit from it. 

In this episode of the Business of Open Source, Campbell discusses the origins of his Company CTO.ai, its revenue model, and why he created this AI companion for developers. He also talks about being a consultant, running a start up, and when it is time to shift gears to building a revenue-focused company.

Highlights:

  • Workflows.sh (4:10)
  • The creations of CTO.ai (5:06)
  • Can being a consultant be a boon for a company (8:48)
  • Revenue models of CTO.ai (11:32)
  • The tipping point of in building a business (15:35)
  • Examples of escape patches (18:20)
  • What is the right amount of money to raise? (20:45)
  • The pricing model of CTO.ai (24:06)
  • Is not having an open-source project limiting? (27:24)

Links:

Kyle

  • LinkedIn
  • Twitter
  • CTO.ai

View Details

Matt Butcher is no stranger to the ways of ethical philosophy. With a Ph.D. in Religion and Computer Science, he enjoys philosophical conversations of ethical dilemmas. Butcher passionately debates wild theories and paradoxical situations against those not afraid to question reality in pursuit of knowledge.

In this episode of The Business of Open Source, Hear how Butcher discusses ethics in the open source world, the grey areas in being an ethical company, and the moral nature of work/life balance. Butcher also details how some of the greatest philosophical minds shaped his own view of ethics and the pursuit of "that middle road."

Highlights:

  • Ethics and Computer Science vs. AI (1:30)
  • Ph.D. in religion and… computer science? (3:05)
  • Ethical claims in open source (4:52)
  • Is it ethical to build a company like yours? (8:50)
  • Interacting with people to see the value (13:45)
  • Understanding balancing creativity with reality(17:52)
  • What does it mean to build an ethical company(19:42)
  • "Complex is such an interesting choice here" (24:43)
  • Finding that middle road (28:29)
  • Common unethical source practices (29:30)
  • Have you changed your ethic since becoming a CEO? (31:50)

Links:

Matt

  • LinkedIn
  • Twitter
  • Fermyon

View Details

Dawn Foster, Director of Open Source Community Strategy at VMware, is a champion of community strategy and development. A doctor of Philosophy, Foster is well-versed in the understanding of collaboration and leverages her mountain of knowledge to fight for the health of maintainers in open-source projects.

In this episode of The Business of Open Source, Dawn Foster joins me from the Open Source Summit North America to tackle community strategy and contribution growth methods. Foster also touches on the differences between open contributions and what project leads should do to help grow their maintainers.

Highlights:

  • Why is it essential to have a contributor growth strategy? (1:46)
  • Loss of control (3:10)
  • How to be proactive for project growth (4:24)
  • Proactive communication to foster a relationship (7:15)
  • Non-code contributions are just as crucial as maintainers (9:47)
  • Is it a mistake to have no contributor growth strategy (12:20)
  • One tactic when being a single maintainer (13:33)
  • Replacing your maintainers (16:02)
  • Don't get arrested (18:10)
  • Improving your skills in maintaining (19:17)
  • Navigating contributions to a project (21:29)
  • Increasing the number of contributions per person (24:13)
  • Example of a good growth strategy (27:07)

Links:

Dawn

  • LinkedIn
  • Twitter
  • VMware

View Details

Bart Farrell is a content creator and community leader in the public speaking world. Based in Spain, he has developed a massively popular platform through podcasting and consulting as a nontechnical person in a technical space.

In this episode, Farrell breaks down the ins and outs of public speaking. He also discusses the art form of speaking offline rather than online, what to expect when talking to an audience, and what it takes to deliver the gift of gab. Farrell also goes into his struggles in becoming a public speaker and his favorite influence on his craft—all this and more in this episode of The Business of Open Source.

Highlights:

  • The Fear of public speaking (2:00)
  • Not a life or death situation (5:10)
  • Be comfortable while speaking(8:40)
  • You only get one start (9:57)
  • A I D A (11:22)
  • Online is not offline (15:03)
  • How to get into public speaking (17:00)
  • Basic building blocks (19:00)
  • How to improve in public speaking (23:35)
  • A favorite public speaker (26:55)

Links:

Bart

  • LinkedIn
  • Twitter
  • bartfarrell.com

View Details

Matt Butcher, CEO of Fermyon, joins me to discuss the ethics of open source and how to keep your company's health in mind when growing your business.

In this episode, Matt and I dig into the ethics of open source and how his background in philosophy influences the decisions he makes as a CEO. We also cover how you can intentionally create and maintain your company values and culture. Finally, Matt reveals his top mistakes as a CEO and how he's overcome them to improve his business.

Highlights:

  • Matt introduces himself and his background in open source and philosophy (0:47)
  • How Matt's background in philosophy changed his perspective on the ethics of open source (3:14)
  • How that background influences how they began and continue to run Fermyon (8:58)
  • How to establish values for a new company and stick to them (15:16)
  • Why Matt started Fermyon when he did and with the focus on web assembly (21:59)
  • Matt reviews the top mistakes he made as a founder and how addressing them has helped him improve the company culture (26:44)

Links:

Matt

  • LinkedIn: https://www.linkedin.com/in/mattbutcher/
  • Twitter: https://twitter.com/technosophos
  • Company: https://www.fermyon.com/
  • Fermyon Discord: https://discord.com/invite/AAFNfS7NGf

View Details

Henrik Rosendahl, CEO of Spiio, joins me to chat about his experience as an entrepreneur and what he’s learned about building successful companies.

In this episode, Henrik and I cover the many aspects of building startups. From the top mistakes new founders make to the best way to monetize your open source business. Listen to learn Henrik’s thoughts on entrepreneurship, including monetization, the three things you need to build a successful startup, and whether or not founding a startup should always feel like a struggle.

Highlights:

  • Henrik introduces himself and gives a brief overview of his experience as an entrepreneur. (0:43)
  • The top mistakes entrepreneurs make building their first startup. (2:25)
  • Monetization and valuable feedback from customers (9:05)
  • How the relationship between the user and the buyer impacts startups. (14:27)
  • Focusing on enterprise vs. a mid-market segment. (16:57)
  • What’s next for Henrik (19:34)

Links:

Henrik

  • LinkedIn: https://www.linkedin.com/in/hrosendahl
  • Twitter: https://twitter.com/hrosendahl
  • Company: https://www.spiio.com/

View Details

Today I’m talking with Fabian Pinckaers, CEO and Founder of Odoo, a suite of business apps to manage all of a business’s activities, about his passion for open source and knowing how and when to pivot as a start-up.

In this episode, Fabien and I discuss the highs and lows of running a start-up as he details his history with Odoo. From its inception as a service offering for auction houses to its current state as an open core software vendor with a cloud offering, Odoo has challenged its founder to continue innovating the product and pivoting the business model to find success. Listen in to discover the lessons Fabien has learned in his journey as a founder and CEO.

Highlights:

  • Fabien introduces himself and explains how his passion for open source led him to start Odoo (0:49)
  • Why Fabien agrees with other founders that open source is a development model, not a business model (5:09)
  • How to keep going forward when your business is struggling. (9:53)
  • Why Fabien believes that “product is everything,” and how that philosophy relates to his passion for open source (10:51)
  • “Open source is about the masses” (14:06)
  • Odoo’s three pivots (19:47)
  • How pivoting to open core has allowed Odoo to contribute even more (24:17)

Links:

Fabien

  • LinkedIn: https://be.linkedin.com/in/fpodoo
  • Twitter: https://twitter.com/fpodoo
  • Company: https://www.odoo.com/

View Details

Today I’m chatting with Nicklas Gellner, co-founder of Medusa, “the open source Shopify alternative” about why he started the company, why open source, and his vision for the future.

In this episode, Nicklas details the original inspiration for Medusa as well as why they chose the name. We also review the switch from an agency to a product focused company. When I mention the buzz around a relatively young company like Medusa, Nicklas emphasizes Medusa’s developer-first approach and explains how that encourages community development.

Highlights:

  • Nicklas introduces himself and Medusa (0:46)
  • Nicklas explains the original inspiration for Medusa’s creation (2:09)
  • How Medusa differs from Shopify beyond being open source and its real value for developers (10:48)
  • How and why Nicklas switched from and agency to a product focused company with Medusa (14:46)
  • The buzz around Medusa (20:44)
  • How Medusa’s developer-first approach encourages community development (22:53)
  • Nicklas’s vision for Medusa’s growth in the coming years. (24:09)
  • How they decided on the name “Medusa” (26:17)

Links:

Nicklas

  • LinkedIn: https://dk.linkedin.com/in/ngellner
  • Twitter: https://twitter.com/nicklasgellner?lang=en
  • Company: https://medusajs.com/
  • Medusa Community Discord: https://discord.com/invite/medusajs

View Details

Michael Chenetz, Head of Technical Marketing at Portainer.io, joins me to discuss why technical marketing is so much more effective for open source companies and also why it’s a hard role to fill.

In this episode, Michael and I discuss his unique background that lead him to technical marketing in the open source space, and the importance of relating tech to people. Michael explains the differences between traditional marketing and technical marketing, as well as the impact technical marketing has on a company’s trajectory.

Highlights:

  • Michael introduces himself and his role at Portainer.io (00:52)
  • The difference in perception and practice for technical marketing versus traditional marketing (02:39)
  • Michael’s path to becoming a technical marketer (05:59)
  • The open source behind Portainer and Michael’s learnings marketing an open source company (08:43)
  • What Michael sees as the relationship between cloud native and open source (13:06)
  • The different impacts technical marketing can have based on company size (18:44)
  • Some of the biggest mistakes founders of open source companies make, according to Michael (21:36)

Links:

Michael

  • LinkedIn: https://www.linkedin.com/in/mchenetz/
  • Twitter: https://twitter.com/mchenetz
  • Company: https://portainer.io

View Details

Liz Rice, Chief Open Source Officer at Isovalent, joins me to discuss the business model behind Cilium and the enjoyment she has found working in open source.

In this episode, Liz and I discuss why Isovalent decided to donate Cilium to CNCF, and the additional decisions behind developing for Cilium open source versus Cilium for Enterprise. Tune into this episode to hear how entrepreneurship taught Liz what she didn’t enjoy doing so she could focus on work she enjoys, and what she finds most rewarding about working in open source.

Highlights:

  • Liz introduces herself and describes her role as Chief Open Source Officer at Isovalent (00:55)
  • Why Isovalent decided to donate Cilium to CNCF (02:11)
  • What Liz sees as the relationship between cloud native and open source (07:43)
  • Liz’s past experiences as an entrepreneur and how it led her to to where she is now (10:05)
  • How Isovalent has evolved and grown into a company with enterprise product offerings (17:17)
  • How decisions are made differently when developing the open-source version of Cilium versus the enterprise version (22:12)
  • The gratification and value Liz has found working in open source (25:58)

Links:

  • LinkedIn: https://www.linkedin.com/in/lizrice
  • Twitter: https://twitter.com/lizrice
  • Github: https://github.com/lizrice
  • Company: https://isovalent.com/
  • Cilium: www.cilium.io
  • eBPF: www.eBPF.io

View Details

Andrew Aitken, Global Open Source Leader at Wipro, joins me to discuss the benefits and challenges that come with adopting open source as a large-scale enterprise organization.

In this episode, Andrew and I discuss the three stages of open source maturity from curiosity to full-scale adoption and mastery. Tune into this episode to learn more about the trigger events that cause large-scale enterprises to explore open source, as well as the barriers and legal challenges they must overcome to effectively adopt it, and most importantly - why Andrew feels it’s so important that every large-scale enterprise goes through this process.

Highlights:

  • Andrew introduces Wipro and explains what it means to be a Global Open Source Leader (00:52)
  • The challenges Andrew helps his clients solve in the world of open source (03:16)
  • How Andrew sees high-security companies like banks view and approach open source (04:28)
  • Do Andrew’s clients view open source software as being more, less, or just as secure as closed source software? (07:41)
  • The three different open-source maturity stages Andrew sees at large companies (10:47)
  • Why Andrew feels most large enterprises would benefit from reaching full maturation in their open-source strategy (14:52)
  • The biggest barriers Andrew sees companies run into when considering moving towards open-source maturity (17:53)
  • Is being enterprise-ready truly a requirement for an open-source start-up? (20:52)
  • The legalities of adopting open source as a large-scale enterprise (25:24)

Links:

  • LinkedIn: https://www.linkedin.com/in/opensourcestrategy/
  • Email: andrew.aitken@wipro.com
  • Company: https://www.wipro.com/

View Details

Randy Abernethy, Managing Director at RX-M, joins me for a chat about the relationship between open source and cloud native.

In this episode, Randy and I discuss how the clients he works with at RX-M are looking to cloud-native as part of their forward-thinking strategies. Tune into this episode to learn how Randy sees the C-Suite viewing open source, how he sees clients evaluate risk in open-source projects, and his views on the relationship between open source and cloud native.

Highlights:

  • Randy introduces himself and his company, RX-M (00:46)
  • Why do companies come to RX-M for help when evaluating and implementing cloud native? (03:44)
  • How do RX-M clients approach open source? (9:35)
  • What are the C-Suite’s views open source (16:19)
  • How Randy’s clients evaluate risk in open-source projects (23:34)
  • The role CNCF plays in how companies evaluate and implement cloud-native solutions (30:31)
  • Randy’s view on the relationship between open source and cloud native (32:39)

Links:

  • LinkedIn: https://www.linkedin.com/in/randyabernethy/
  • Twitter: @randyabernethy
  • Company: www.rx-m.com

View Details

Ian Tien, CEO and Co-Founder of Mattermost, joins me to talk about how Mattermost went from being a video game company to an open-source messaging platform that provides collaboration for developers and other mission-critical teams.

In this episode, Ian and I discuss the reality of product-led growth in open-source companies, Ian’s perspective on open source moving towards platform-based solutions, and the advice he would give to other open-source founders.

Highlights:

  • Ian introduces himself and the Mattermost open-source project (00:45)
  • The use cases Ian sees for Mattermost (05:17)
  • Ian takes us through the origin story of Mattermost and how it went from being a video game company to an open-source messaging solution (08:59)
  • The role open source played in the success of Mattermost (15:07)
  • Ian’s perspective on open source moving towards platform based solutions (20:22)
  • Does Ian think the product-led growth model of “If you build it, they will come” is realistic, and how can that mentality lead to success? (27:34)
  • The advice Ian would give other open-source founders (28:33)

Links:

Ian

  • Twitter: https://www.linkedin.com/in/iantien/
  • Company: https://www.mattermost.com

View Details

Amanda Brock, CEO of Open UK, joins me for an engaging conversation on best practices in founding an open-source company.

In this episode, Amanda and I chat about the various business models available for building a company around open-source technology, the common pitfalls and crossroads open-source founders find themselves facing, and how to do open-source in a way that leads to long-term success and profitability.

Highlights:

  • What is Open UK? (00:40)
  • The various business models for building a company around open-source technology (04:09)
  • Which business models Amanda feels work best and why (08:07)
  • The importance of founders prioritizing open-source communities (14:07)
  • How and why to do open-source the right way (17:04)
  • What is the true cost of founding an open-source company compared to traditional business models? (26:44)
  • Who are you building for, and how do you get to profitability? (30:35)

Links:

  • LinkedIn: https://www.linkedin.com/in/amandabrocktech/
  • Twitter: www.twitter.com/amandabrocktech
  • Open UK www.twitter.com/openuk_uk
  • Company: www.openuk.uk

View Details

Wes McKinney, CTO & Co-Founder of Voltron Data, joins me for an in-depth conversation on how his quest to develop Python as an open-source programming language led him to creating the pandas project and founding four companies.

In this episode, Wes and I dive into his unique background as the founder of the pandas project and he describes his perspective on the early days of Python, his journey into the world of open-source start-ups, and the risks and benefits of paying developers to work on open-source projects.

Highlights:

  • Wes introduces himself and describes his role (00:46)
  • Wes’ role in elevating Python to a mainstream programming language (02:15)
  • How working with Python led Wes to co-founding his first two companies (09:01)
  • Apache Arrow’s critical role at Voltron Data and their focus on accelerating Arrow adoption (12:52)
  • How did the team at Voltron Data decide on an open-source business model? (18:54)
  • Wes speaks to the risk that can come from having developers work on an open-source project (22:31)
  • Wes’ perspective on the real-world applications and benefits of paying developers to work on open-source projects (27:44)

Links:

Wes

  • LinkedIn: https://www.linkedin.com/in/wesmckinn/
  • Twitter: https://twitter.com/wesmckinn
  • Company: https://voltrondata.com/

View Details

Braden Hancock, Co-founder and Head of Technology at Snorkel AI, joins me to talk about his path from academia to start-up co-founder and his vision to make AI more accessible to both traditional and no-code development.

In this episode, Braden and I explore the journey he and his co-founders took to go from having an interesting idea to forming a company and the strategic business decisions they made along the way, such as why they opted not to use an open-source business model and the educational marketing strategy they’ve adopted.

Highlights:

  • Braden discusses his role as co-founder of Snorkel AI. (00:25)
  • An introduction to Snorkel Flow, Snorkel AI’s data-centric AI development program and the challenges they solve for. (01:49)
  • Snorkel AI’s relationship with open source. (06:30)
  • Why Snorkel AI decided not to use an open-source business model in order to lower the barrier to entry. (09:01)
  • Snorkel AI’s trajectory coming from academia to the world of start-ups. (12:50)
  • The unexpected challenges of building Snorkel AI. (17:50)
  • Taking an educational approach to the marketing at Snorkel AI. (22:27)
  • Braden discusses the meaningful applications of AI as well as where he sees AI being used as more of a buzzword. (27:27)

Links:

Braden

  • LinkedIn: https://www.linkedin.com/in/bradenhancock/
  • Twitter: https://twitter.com/bradenjhancock
  • Company: snorkel.ai
  • Snorkel AI Twitter: https://twitter.com/SnorkelAI

View Details

Anna Filippova, Director of Community & Data at DBT Labs, joins me to chat about the fundamental role community plays in the world of open source and her role helping to create a thriving community.

In this episode, Anna and I dive into the concept of a community: why it’s essential for open-source development, how to create business value through community, and how to track community health above and beyond user count.

Highlights:

  • Why community is mission critical at DBT Labs (02:07)
  • The fundamental role open source played in creating DBT Labs as its known today (05:57)
  • The approach DBT Labs uses to create business value through community (08:13)
  • Anna’s framework for the three buckets of communities (09:42)
  • Why measuring and tracking community health is a more valuable metric than just user count (11:34)
  • What do people get out of communities, and why are they so valuable? (19:06)
  • Common misconceptions around building communities as a business strategy (24:19)

Links:

  • Anna
    • LinkedIn: https://www.linkedin.com/in/annafilippova/
    • Twitter: https://twitter.com/anna_fil
  • DBT Labs: getdbt.com/community
  • Coalesce Conference: https://coalesce.getdbt.com/

View Details

Highlights:

  • Open source software at the department of defense (1:36)
  • Is there risk associated with using open source software in the department of defense? (5:30)
  • Does the public sector contribute to and participate in open source communities? (9:13)
  • Rob’s background and work experience (14:25)
  • What led Rob to found Defense Unicorns (16:35)
  • Rob’s focus on a specific niche in the founding of his company (17:33)
  • How working with a fixed budget affects an open source company (19:33)

Links:

Rob

  • LinkedIn: https://www.linkedin.com/in/robertcslaughter/
  • Company: https://www.defenseunicorns.com/

View Details

Ricardo Mendoza, founder and CEO of Pantacor, joins me for a chat at the Open Source Summit in Austin. Ricardo shares why he started Pantacor and describes the differences between IoT, edge, connected, and embedded devices. I ask him how Pantacor fits into the edge continuum, and he explains how Pantacor helps bring embedded devices into the future. Ricardo talks about the open source arm of Pantacor’s strategy, we discuss Pantacor’s unique interest in hardware versus primarily dealing with software, and Ricardo wraps up by sharing his advice for aspiring business owners!

Highlights:

  • Why Ricardo started Pantacor (1:19)
  • Difference between IOT edge devices, connected devices, and embedded devices (2:17)
  • How Pantacor fits into the edge continuum (4:49)
  • Why are embedded systems lagging behind and how does that manifest? (6:22)
  • How open source is part of Pantacor’s strategy (9:40)
  • How aware are manufacturers of their operating systems and how Pantacor could help them? (13:35)
  • Pantacor’s relationship with hardware (16:45)
  • What was the inspiration for the founding of Pantacor? (20:11)
  • The difference between cloud developers and their relationship with open source versus the relationship between embedded devices and open source (22:46)
  • Is there a disadvantage to being based in Europe? (24:51)
  • Advice for someone who wants to start a company or work with embedded devices (26:28)

Links:

Pantacor

  • https://pantacor.com/
  • https://pantavisor.io/
  • Twitter: @pantahub

View Details

Live from the Open Source Summit in Austin, I sit down with Jeff Shapiro, the License Scanning Manager for the Linux Foundation. Jeff begins by explaining what he does at the Linux Foundation, including ensuring that open source licenses are compatible and compliant. We discuss what license issues start-ups should be aware of, how to educate yourself on open source licensing, and when you should consult an expert. Jeff clarifies some confusion around dual licenses and explains the challenges of changing licenses on an open source project. Finally, we discuss the possibilities of disallowing specific uses through licensing and who can write a license.

Highlights:

  • Jeff talks about the legal and business risks of non-compliant open source licenses (3:09)
  • License issues start-ups should be aware of (7:16)
  • DCO (Developer certificate of origin) and understanding where code comes from (12:10)
  • Educating yourself and others about open source licenses (13:04)
  • Jeff talks about when you need to consult an expert (15:36)
  • Jeff explains how he got into licensing as an engineer (17:23)
  • Jeff discusses dual licenses (18:18)
  • How hard is it to change licenses on an open source project (20:23)
  • Jeff explains if it’s possible to disallow specific uses with your license (23:39)

Links:

Jeff

  • LinkedIn: https://www.linkedin.com/in/jeffcshapiro/
  • Company: https://www.linuxfoundation.org/

View Details

Today I sit down with Webb Brown, CEO and cofounder of Kubecost. Kubecost provides real-time cost visibility and insights for teams using Kubernetes. Webb tells the story of building Kubecost, starting with the pain points that inspired the open source tool. He talks about the transition from an open source project to becoming a commercial company, and explains the decision to build a company with the same name and branding as the open source tool. Webb talks about Kubecost’s newest initiative, OpenCost, and concludes by offering some lessons and advice for anyone in the early days of an open source startup.

Highlights:

  • Webb explains what Kubernetes cost is (1:27)
  • How the pain points addressed by Kubecost usually manifest (3:04)
  • What the impetus was for building the Kubecost open source tool (5:30)
  • The transition from open source to commercial (6:54)
  • The relationship between a cost-cutting tool and open source (10:48)
  • Kubecost’s new initiative, OpenCost (13:40)
  • The decision to have a company with the same name as the open source project (18:55)
  • Pros and cons that are unique to building an open source company (22:08)
  • Advice for anyone in the early stages of an open source startup (25:22)

Links:

Webb

  • Email: webb@kubecost.com
  • LinkedIn: https://www.linkedin.com/in/webbbrown/
  • Twitter: https://twitter.com/webb_brown
  • Company: https://www.kubecost.com/

View Details

Today I sit down with Ev Kontsevoy, the CEO and co-founder of Teleport, a software company that began as an open source project. Teleport is an identity aware multi protocol access proxy that Ev was inspired to create because of the inherent frustrations with security he experienced in his career. Ev talks about how Teleport began as an open source tool and then grew into enterprise. I ask Ev what things he has done differently from his first start-up, Gravity, and we discuss how the open source community culture has bled into the company culture at Teleport. We end by talking about the SaaS version of Teleport and the ways in which the open source version funnels business into the commercial version.

Highlights:

  • Security frustrations that led to the founding of Teleport (1:17)
  • Ev talks about Teleport’s vision and how it began as an open source project (6:33)
  • Ev talks about Teleport’s first customer and a separate open source project, Gravity (12:09)
  • How Ev’s experience with a prior start-up changed his approach to Teleport (18:24)
  • Ev discusses the culture and community at Teleport (21:16)
  • How Teleport chooses which features to keep open source and which ones to offer as commercial (24:38)
  • The SaaS version of Teleport (26:55)
  • The different audiences for the different iterations of Teleport (28:08)

Links:

Ev Kontsevoy

  • LinkedIn: https://www.linkedin.com/in/kontsevoy/
  • Twitter: https://twitter.com/kontsevoy
  • Company: goteleport.com

View Details

Today I’m joined by Cillian Kieran, the CEO and co-founder of Ethyca, to talk about the privacy challenges that served as the impetus to found Ethyca. In our chat, he explains the overarching goals of the privacy engineering platform. We discuss the decision to begin Ethyca as an open source tool and why that was critical to the mission. Then we talk about the decision to move to a commercial product and how to decide which features to offer as paid versus free. Cillian reviews the differences in his process between his two start-ups, discusses lessons he learned from prior mistakes, and provides advice for aspiring founders of open source start-ups.

Highlights:

  • How Cillian decided to found Ethyca (00:50)
  • Awareness of developers and engineers around privacy issues (3:46)
  • Cillian talks about why he went the open source route (8:15)
  • Moving from open source to commercial product (14:02)
  • Privacy as a human right and how that influences development of features (16:32)
  • How Ethyca manages relationships between engineer and legal teams (19:40)
  • What Cillian did differently at his two start-ups (21:58)
  • We discuss open source start-up success and whether it’s necessary to have a larger world-changing vision (24:52)
  • Cillian discusses mistakes he has learned from (27:56)
  • Cillian offers advice to aspiring founders in the open source community (30:49)

Links:

  • Fides open source platform: fid.es

Cillian

  • Twitter: @Cillian
  • Company: https://ethyca.com/

View Details

Today I’m joined by CEO and founder of LeanIX, André Christ. André begins by describing his business, and then explains how his experiences working in large enterprise inspired him to build a product that would help businesses catalogue their software and optimize their portfolios. André offers advice for companies desiring to sell primarily to enterprise and expounds on the his experience with the differences between traditional enterprise and large enterprise. We discuss LeanIX’s transition to become a global company based in Europe, and conclude our talk with some advice from André to potential founders.

Highlights:

  • André describes his his company LeanIX (00:48)
  • The experiences that led André to found LeanIX (2:50)
  • LeanIX’s decision to focus on enterprise customers (7:45)
  • Advice for companies that want to focus on selling to enterprise (9:47)
  • The difference between traditional enterprise and very large enterprise like Amazon (15:19)
  • Transitioning to becoming a global company based in Europe (19:49)
  • The surprisingly fragmented world of global tech (24:50)
  • LeanIX’s decision to expand into other products (26:30)
  • André’s advice for anyone considering starting a company (27:57)
  • André shares about scaling mistakes and how LeanIX has learned from them (31:48)

Links:

André

  • LinkedIn: https://www.linkedin.com/in/andrechrist/
  • Twitter: https://twitter.com/christ_andre
  • Company: https://www.leanix.net/

View Details

Today I chat with Keith Basil, GM of Edge Computing at SUSE. We begin by reviewing the definition of edge: Keith explains how SUSE breaks edge computing down into 3 categories, and then talks about the shared understanding of edge by the industry at large. I ask Keith about the overlap of edge products with non-edge products, and then we discuss the maturity of the edge landscape and Keith explains how SUSE helps clients with infrastructure. We wrap up by talking about managing feature bloat and SUSE’s decision to have their entire code base be open sourced.

Highlights:

  • Keith breaks down the 3 categories of “edge” as defined at SUSE (1:14)
  • We discuss the industry understanding of edge technology (5:34)
  • Keith defines “edge watching” (8:44)
  • We discuss the relationship between cloud native and edge native (10:22)
  • The overlap of edge products and non-edge products (14:25)
  • The maturity of the edge landscape and how SUSE help clients with infrastructure (17:04)
  • How SUSE manages feature bloat (23:37)
  • SUSE’s decision to have their entire code base be open sourced (26:15)

Links:

Keith

  • Twitter: @noslzzp
  • Company: https://www.suse.com/

View Details

Today I talk with Shaun O’Meara, the global field CTO at Mirantis. We begin by discussing the integration of Docker Enterprises with Mirantis approximately three years ago. We discuss the challenges of integrating companies, including incorporating new technology, processes, and customers and merging two very different work cultures. Shaun offers his advice for anyone considering selling to enterprises and emphasizes the role of partnering with customers and becoming part of their process. Shaun talks about the expectations and realities of merging Docker and Mirantis, including the challenges of a licensing model change. We conclude our time by discussing the differences between selling to small companies versus selling to enterprises.

Highlights:

  • How integrating Docker Enterprises with Mirantis affected Shaun’s role as CTO (1:09)
  • How incorporating Docker technology helped Mirantis build different value for customers (3:39)
  • We talk about the effects of combining the work cultures of Docker and Mirantis (5:40)
  • Shaun offers advice for people considering start ups or selling to enterprise, including the importance of partnering with customers (8:47)
  • Shaun talks about his expectations of merging Docker and Mirantis versus reality (12:56)
  • We talk about the licensing model change through the transition (14:34)
  • Shaun talks about outsourcing versus what Mirantis does in augmenting and supporting teams (17:55)
  • We discuss the differences between selling to small companies and enterprise (20:06)

Links:

Shaun

  • LinkedIn: https://www.linkedin.com/in/shaun-omeara/
  • Company: https://www.mirantis.com/

View Details

Today I sit down and chat with John McBride, senior software engineer at VMware. We begin by talking about John’s address at KubeCon, “Risks of Single Maintainer Dependencies and How to Mitigate Those Risks.” We discuss the definition of security and then John identifies some of the other non-security risks posed by single maintainer dependency. We talk a little bit about mitigating the risks and about building trust and community around single maintainer projects. We conclude our time by speculating on the extinction of single maintainer dependencies.

Highlights:

  • John introduces himself and talks about his interest in mitigating the risks of single maintainer dependencies (00:55)
  • We have a conversation about the definition of security (4:54)
  • John talks about the other, non-security risks of single maintainer dependency (10:00)
  • We discuss how to mitigate the risks of single maintainer dependency (12:04)
  • John talks about building trust and building community around single maintainer projects (16:48)
  • John answers my question “Do you think being a single maintainer is ultimately an anti-pattern, a non best practice?” (23:56)

Links:

John

  • Twitter: @johncodezzz
  • Company: https://www.vmware.com

View Details

Today I talk with Catherine Paganini, Head of Marketing and Community at Buoyant. We begin by discussing the Cloud Native Glossary and how it is helping to make cloud native concepts more accessible for people around the world. Catherine talks about nurturing community in open source projects, and about the function of documentation. Catherine and I discuss pitfalls in building open source communities, and Catherine talks about her strategy for recovering from mistakes. Catherine concludes the conversation by talking about balancing her roles as head of marketing and community at Buoyant.

Highlights:

  • Catherine talks about how the Cloud Native Glossary started, how it has grown, and how it helps to make education about the cloud accessible and easy to understand (1:00)
  • Catherine discusses about how the Cloud Native Glossary is being used (5:47)
  • Catherine and I talk about nurturing community in an open source project (8:22)
  • Catherine discusses empowering end users through efforts like the Linkerd Anchor Program (11:28)
  • Catherine talks about the function of documentation (14:05)
  • I ask Catherine, “What do you see people getting wrong when it comes to nurturing community?” (15:29)
  • Catherine talks about recovering from mistakes (18:49)
  • Catherine discusses walking the line between being head of marketing and head of community (23:47)

Links:

Cloud Native Glossary: https://glossary.cncf.io/

Linkerd: https://linkerd.io/

Linkerd Anchor Program: https://linkerd.io/community/anchor/

Catherine

  • LinkedIn: Catherine Paganini
  • Twitter: @cathpaga
  • Company: https://www.buoyant.io

View Details

Today I talk with Yann Léger, CEO of Koyeb, the serverless developer platform that allows businesses to safely and easily deploy applications. We begin by talking about Yann’s decision to base the company on serverless, and the true meaning of cloud native. Yann then discusses Koyeb’s relationship with Kuma, and Koyeb’s posture towards open source projects. The conversation concludes with Yann sharing mistakes he’s learned from in the process of building Kuyeb and offering advice to other potential technical founders.

Highlights:

  • Yann talks about the decision to leave his position at Outscale and start his own company (1:44)
  • Yann discusses choosing to base his company on serverless (3:14)
  • Emily and Yann talk about the meaning of cloud native (6:00)
  • Yann talks about Koyeb’s relationship with Kuma (9:40)
  • Yann discusses Koyeb’s open source projects (11:46)
  • Yann shares mistakes he has learned from in the process of building Koyeb (15:25)
  • Yann answers the question “What are the disadvantages of being a technical founder?” (18:06)
  • Emily and Yann discuss the challenges of remote working (22:00)
  • Yann’s advice for anyone considering becoming a technical founder (23:15)

Links:

Yann

  • LinkedIn: Yann Léger
  • Twitter: @yann_eu, @gokoyeb
  • Company: https://www.koyeb.com/

View Details

Today I chatted with Dirk Hohndel, chief open source officer at the Cardano foundation. We begin by defining an open source ecosystem, and then talk about what different open source ecosystems might look like and how they are maintained. Dirk talks about best practices for steering an open source ecosystem, and then we discuss the role of foundations in open source projects. I ask “how do you define success for an open source project” and we end with a discussion on the best practices for running open source project foundations.

Highlights:

  • We talk about the meaning and maintenance of an open source ecosystem (1:31)
  • The differences between an open source ecosystem and a community (11:06)
  • Dirk talks about best practices for steering an open source ecosystem (13:00)
  • The role of a foundation in an open source project (18:04)
  • Dirk discusses other iterations of open source projects that can be successful (22:19)
  • Dirk answers the question “how do you define success for an open source project?” (24:43)
  • We discuss best practices for running an open source project foundation (27:49)

Links:

Dirk

  • Twitter: @_dirkh
  • Company: cardanofoundation.org

View Details

Today I sit down with Romaric Philogène, CEO and founder of Qovery, a platform that helps developers build, deploy, run, and scale applications. Romaric begins by talking about his first two start-ups, both social networks, and then we discuss the difference between creating consumer-facing products and products for developers. We then talk about marketing in the US as it compares to the global market. We discuss Qovary’s relationship to open source and the idea of fostering community around a company’s culture. Romaric concludes by offering advice to developers on the value of being a skilled communicator.

Full Description / Show Notes

Highlights:

  • Romaric talks about his first two startups that preceded Qovery (3:26)
  • The differences between building a consumer facing product and creating a product for developers (5:55)
  • Romaric talks marketing in the US vs marketing in Europe (11:50)
  • Romaric answers the question “what are things you’re doing differently now that you’ve learned from previous efforts?” (17:00)
  • The value of community building in marketing to developers (19:28)
  • Qovary’s relationship with open source (20:46)
  • Building community around your company vs just a product (25:00)
  • The importance of communication as an engineer (28:47)

Links:

Romaric

  • Twitter:@rophilogene
  • Company: https://www.qovery.com/

View Details

Today I’m chatting with Avery Pennarun, CEO of Tailscale. Tailscale is a VPN service that makes devices and applications accessible anywhere in the world by enabling encrypted point-to-point connections using the open source WireGuard protocol. Avery begins by talking about his experience building a start-up while he was a college student and how things have changed as he leads his current start-up. Avery recommends the book “Crossing the Chasm” and we discuss market segmentation as it relates to creating a successful start-up. Avery explains how Tailscale has been successful in implementing market segmentation strategies. We conclude our conversation by talking about goal setting and the importance of quality.

Highlights:

  • Avery talks about his first start-up experience as a college student (1:10)
  • Avery recommends “Crossing the Chasm” and discusses how it influenced his start-ups (7:54)
  • We discuss market segmentation strategy (13:29)
  • Specific marketing strategies used at Tailscale (18:41)
  • Avery talks about mistakes he’s made while building his start-ups (22:24)
  • Goal setting in start-ups (24:42)
  • We talk about the importance of quality in building word of mouth success (29:49)
  • Avery answers the question “How do you maintain an identity as an engineer when you are also a serial entrepreneur?” (33:23)

Links:

Avery

  • Twitter: @apenwarr @tailscale
  • Company: https://www.tailscale.com

View Details

Today I chat with Kiersten Gaffney, CMO of Codefresh, a software delivery platform. Kiersten begins by defining her role as CMO. We then discuss the unique challenges of product strategy with open source projects. Kiersten talks about the importance of maintaining both a top-down and bottom-up approach when taking a project from open source to enterprise, and then explains some of the most common mistakes she’s seen when companies undergo this process. We discuss how technical a team should be when marketing open source and conclude the conversation by talking about analysis paralysis in start-ups and how to avoid it.

Highlights:

  • Kiersten answers the question “What do CMOs do all day?” (1:49)
  • Product strategy with open source products (2:58)
  • How open source projects fit into marketing efforts (6:10)
  • Kiersten’s advice on how to maintain both a top-down and bottom-up approach (11:28)
  • Is there a magic formula for taking a project from open source to for profit? (13:02)
  • Biggest mistakes when taking a project from open source to enterprise (15:54)
  • Emily asks how technical a marketing team should be for an open source project (22:53)
  • Kiersten and Emily discuss the tension between engineering and marketing (24:24)
  • Analysis paralysis in startups (26:35)

Links:

Kiersten

  • LinkedIn: https://www.linkedin.com/in/kierstengaffney/
  • Twitter: @kierstengaffney
  • Company: https://codefresh.io/

View Details

In this short summation episode, I talked a little more about why I think Deepfence's open source strategy is so genius. 

View Details

Today I sat down with Sandeep Lahane, founder and CEO of Deepfence, a security preventive and detective solution for cloud and container native environments. Sandeep began by explaining both the open source and enterprise components of Deepfence. Threatmapper is a multi-cloud platform for scanning, mapping, and ranking vulnerabilities in running containers, images, hosts, and repositories, and Threatstryker is a commercial product that offers runtime attack analysis, threat assessment, and targeted protection for infrastructures and applications. We then talk about the inexhaustibility and the ever-changing landscape of cybersecurity. Sandeep explains the impetus for launching Deepfence and the process of getting to Threatmapper and Threatstryker, and then talks about his journey from working as a systems programmer to launching a tech startup. We discuss the tense relationship between security and development in the industry, and end the conversation with some words of advice for engineers considering the entrepreneurial plunge.

Highlights:

  • What is the difference between Threatmapper and Threatstryker at Deepfence? (00:55)
  • Sandeep explains how the Deepfence commission product builds upon the open source one (2:14)
  • Discussing the inexhaustibility of the cybersecurity landscape (11:40)
  • The genesis of Deepfence (13:58)
  • Sandeep discusses the business benefits of having an open source project (14:57)
  • Sandeep talks about his journey from systems programmer to tech startup (17:20)
  • Emily and Sandeep discuss the tense relationship between security and development (21:19)
  • Sandeep gives advice to engineers considering entrepreneurship (33:57)

Links:

Sandeep

  • https://deepfence.io

View Details

Some short thoughts on marketing in the open source ecosystem, drawn from my conversation with Eric on Wednesday. 

View Details

Today I sit down with Eric Hendricks, the technical marketing director at Red Hat. Red Hat delivers open source solutions that make it easier for enterprises to work across platforms and environments. Eric begins the conversation by discussing his start as a technologist and how he decided to make the move to marketing. Eric then discusses the challenges of bringing marketing savvy into the devops space, including the unintended consequences of marketing buzzwords. I ask Eric about the relationship between marketing and open source, and Eric talks about how many of Red Hat’s community marketing efforts are driven through upstream communities. We then discuss the concept of the buyer in open source versus start ups, and how the difference is that the “big ask” in open source projects is emotional investment. Eric concludes the conversation by talking about the impact of his current role as a technical marketer as compared to the impact of a founder or IC.

Highlights:

  • Eric answers the question “At what point did you start to see yourself and did other people see you as a marketer?” (5:53)
  • The stigma around marketing and the problem with marketing buzzwords in devops (10:22)
  • Emily and Eric discuss the shared vocabulary problem that can arise with newer concepts in tech (14:35)
  • Eric talks about equipping his product’s “champions” with all the resources they need to communicate need and efficacy to potential buyers (15:39)
  • Emily asks Eric about the relationship between marketing and open source (19:50)
  • Emily and Eric discuss the concept of “the buyer” with open source (25:14)
  • Eric answers the question: “how are you able to have more of an impact in your current role than you would as an IC?” (29:30)

Links:

  • Personal website: www.itguyeric.com
  • Company website: www.redhat.com
  • Twitter: @itguyeric
  • Company: @rhel
  • Podcasts: RHEL Presents, Into the Terminal

View Details

I'm trying something new this week — adding an extra episode with some key takeaways from the interview earlier in the week. In this one, I talked about the education battle many cloud native companies face, the problem of open source projects that are too good and understanding pain points for different personas. 

View Details

Today I’m joined by Omri Gazitt, founder of Aserto, an authorization company that offers cloud native authorization as a service. We discuss the differences between ID authorization and authentication and the problems associated with educating developers on the distinctions. Omri also talks about the evolution of authorization from server software all the way to cloud native authorization. He then expounds on the strategic nature of the decision to open source or not, and offers advice to developers based on his experience as both an engineer and an executive.

Highlights:

  • The beginnings of Aserto (7:50)
  • Omri talks about what it was like to be part of a startup, Neon, in the 90s (10:49)
  • Emily and Omri discuss what authorization was like pre-cloud native (13:32)
  • How integration became the strategy used by Aserto to begin to solve authorization problems (17:10)
  • The decision to open source and how organizations should be strategic when considering open source (18:55)
  • Omri discusses his unique perspective as both a former tech engineer and executive in forming his start-up (25:08)
  • Omri talks about a missed opportunity in the early stages of Aserto (28:56)

Links:

Omri

  • LinkedIn:
  • Twitter
  • Company

View Details

I’m joined by Mark Grover, one of the founders of Stemma, a data catalogue for building decentralized data informed cultures. In essence it is a “Google Search” built for data scientists, data analyst, business leaders, and more. Stemma is striving to solve data documentation and relevance issues by keeping data cataloguing up to date and current.

In this episode Mark covers his transition from the data team at Lyft to establishing Stemma. He discusses how he identified the need for a source of truth for ETA data, and how the data scientist in these teams should be the end all for this knowledge. Starting with building Amundsen, Stemma expands on the groundwork laid there to bring data to the user and open-source community’s needs.

Highlights:

  • An introduction to Mark and Stemma (00:00)
  • The history of Stemma (00:45)
  • How open source helped solved data cataloguing problems (4:40)
  • The decision to found Stemma (07:40)
  • How Stemma’s relationship with Amundsen has evolved (13:20)
  • The unexpected challenges and unexpected eases (18:35)
  • Navigating the co-founding experience (23:35)
  • Mark’s vision for Stemma’s future (28:54)
  • Mark’s tips for aspiring founders (32:54)

Links:

Mark

  • LinkedIn
  • Twitter
  • Company

View Details

I’m joined by Matt Leray, co-founder and CTO Speedscale, an API testing product that applies “real world” stresses to “collect and replay traffic without scripting, simulate load and chaos, and measure performance.” With a history steeped in various aspects of tech, and with time spent in the startup space and cloud native space, Matt brings to the table some encompassing perspectives.

Matt’s career has carried him from monitoring satallite earth stations, to fiber optics, and more recently into cloud native. Matt began in startups, then went to larger companies, then back to startups, which he offers some insight on. Matt has a lot of wisdom to share on entrepreneurship, how the startup space has changed, and how to best navigate that. Matt discusses how Speedscale works as an “traffic replay” platform for APIs and his role there both technically and as a co-founder. Check out the conversation for a list of startup how-tos!

Highlights:

  • Introduction to Matt and Speedscale (00:00)
  • When Matt decided to become an entrepreneur (02:10)
  • Deciding to jump into startups (06:00)
  • What has stayed the same, and what has changed for Matt’s entrepreneurship (10:00)
  • Being selective (14:00)
  • Nailing down the timing and finding the right moment for Speedscale (20:26)
  • Matt’s most controversial view about the cloud native startup space (26:25)
  • Matt’s final thoughts (30:18)

Links:

Matt

  • LinkedIn
  • Twitter
  • Company

View Details

Kit Merker, co-founder and COO at Nobl9, a software reliability platform. Through software-defined SLO’s Nobl9 helps developers, DevOps and reliability engineers deliver more reliable features faster. Kit has had a storied career in tech, and as a result is a great source of wisdom and know how. Especially in regard to navigating the various sides of any given business.

In this episode, Kit offers up some anecdotes from his long history in the software space, and how he transitioned from engieenering to the “business side” of things. He tears down some stereotypical misrepresentations of both sides, and expands on how empathy helps alleviate many of these issues. Kit discusses his partnership experiences, work in M&A, building a “coalition” in open space, and more! Tune in for our conversation for Kit’s emphatic and valuable insight.

Highlights:

  • Introduction to Kit and Nobl9 (00:00)
  • Why Kit decided to transition to the “business side” (03:55)
  • Kit’s reflections on partnerships (06:30)
  • The importance of building a coalition around Kubernetes (10:00)
  • Alleviating developer burnout (18:15)
  • Nobl9 and how it came to be and how they work (20:15)
  • The challenges of being a consultant first (27:00)
  • Recognizing the margins (33:30)

Links:

Kit

  • LinkedIn
  • Twitter
  • Nobl9

View Details

Michael Guarino, founder of Plural, a unified platform for open-source management platform entirely hosted on Kubernetes which creates a fully functional ecosystem around deploying Airflow. Though working in Kubernetes and more, Plural can be used across a wide spectrum of open source projects. Many of which Plural is specifically targeting to make their product appealing to users.

In this episode, Michael talks about how Plural works within the open-source space, and how in using it with Airflow they’ve helped to elminate much of the work needed there. Michael lays out how using Plural makes using Airflow easier on the user, versus taking a DYI approach. Michael discusses avoiding lock-in, the various open source tools they use, working through the early days in COVID, the history of building Plural, and more!

Highlights:

  • Introduction to Michael and Plural (00:00)
  • The open source projects that work with Plural and its advantages (01:30)
  • How using Plural is easier than DYI and avoiding lock-in (04:06)
  • How Plural came to be (08:26)
  • The unexpected difficulties, and the unexpected ease (15:06)
  • Plural and open-source (18:40)
  • Navigating potential roadblocks to community building (23:09)
  • Monetization (26:20)
  • Michael’s thoughts on the future of open-source (28:11)

Links:

Michael

  • LinkedIn
  • Plural

View Details

Today’s guest is Michael Tanenbaum, CEO and co-founder of Mycelial, an edge data platform for distributed local first applications that is built with developers in mind. Myceclial takes the accomplishments of the cloud native movement to bring data that exists outside the data center into the hands of the developers themselves. With a focus on data from the “edge”, which Michael defines as anything from a smart thermostat, to a 5g tower, applications, and more.

In this episode Michael lays out how he and his partners captilized on the oppurtunity of recent innovations in cloud native, and in turn commercialize the need to “get applications out of the data center to work harmoniously with applications in the data center.” He and his co-founders are striving to build complex edge native applications and local native data. Michael breaks down the “three pillars” of edge native to provide some crucial definitions, how he identified the needs Mycelial addresses, the diverse range of obstacles they’ve already surmounted, and more!

Highlights:

  • An introduction to Michael and Mycelial (00:00)
  • When Michael recognized the need for a product like Mycelial (02:32)
  • The “Three Pillars” of edge native (05:11)
  • Disovering an “unsexy” problem and deciding to solve it (09:00)
  • The unforseen difficulties of Mycelial (15:15)
  • The unforseen easy parts of Mycelial (20:32)
  • Some important takaways from the founding experience (26:45)

Links:

Michael

  • Twitter
  • Mycelial

View Details

Lyon Wong, CEO and co-founder of Blameless, a complete reliability engineering platform that brings together AI-driven incident resolution, blameless retrospectives, SLOs/Error Budgets, and reliability insights reports and dashboards that enable businesses to optimize reliability and innovation. Lyon has a history steeped in founding and investing in start ups and company building, which has lead a heavy involvement in Blameless where he can apply the many lessons learned.

In this episode, Lyon breaks down his background and how it influenced his decision to become a founder at Blameless. Over the course of his career he noticed trends in other companies where teams were prevented from learning opportunities because they were worried about catching the blame. As a result Lyon identified the need in the market for a way to synthesize the cultural tensions around blame. Lyon’s insight on building trust, partnership, and communications on learning are deep and worthwhile. Check out the full conversation!

Highlights:

  • An introduction to Lyon and Blameless (00:00)
  • Jumping back into being a founder after time as a VC (2:28)
  • Creating a blameless culture (05:50)
  • What Lyon does different as a founder and investor and his early experiences (09:50)
  • The importance of credibility (16:10)
  • The “core skillsets” needed in start ups and some crucial beliefs (18:55)
  • The larger and smaller pictures, and balancing short and long term (25:34)
  • Lyon’s parting words and wisdom for founders (32:51)

View Details

Haseeb Budhani, co-founder and CEO of Rafay Systems, a Kubernetes operations company joins the conversation. What a Kubernetes operation company is as companies use Kubernetes across their organizations they need the right automation, security, visibility and more. These are needs that come from multiple teams working across multiple applications, and it creates a lot of work. This is where Rafay Systems is looking to cover down.

Haseeb introduces us to the work at Rafay systems, and his own discovery of the problems they want to address. Haseeb discusses the history of Rafay’s establishment, and how they are striving to create a fluid and robust workflow engine. He reflects on how his previous experience has reinforced the lessons he brought to Rafay, how to connect to the customers, and more! Check out the conversation!

Highlights:

  • An introduction to Haseeb and Rafay Systems (00:00)
  • Lessons learned at other companies and staying the course (6:12)
  • Successfully connecting with the customers needs (12:29)
  • The lessons learned already at Rafay and some helpful advice (15:36)
  • Where the Kubernete’s ecosystem is headed (25:22)

Links:

Haseeb

  • LinkedIn
  • Twitter
  • Company

View Details

Andrew Rynhard, Founder and CTO of Sidero Labs, joins the show today to discuss his work and Sidero Labs. Sidero is a Kubernetes lifecycle management reimagined from the operating system to entire stack. Andrew has origins steeped deeply in open source, and it has become a central focus to his entire ethos and drive.

Andrew breaks down his own trajectory that lead to Sidero and the passion he leveled for the endevour from the onset. Andrew’s passion for open source served as the impetus founding the company, and he shares his love for open source and the pathways that it created for him through his career. Andrew shares Sidero Lab’s successful initial funding, the shift to it being his full-time job, and their meteoric rise.

Highlights:

  • An introduction to Andrew and Sidero Labs (00:00)
  • The early days of Sidero Labs and their current position (03:10)
  • The moment Sidero became Andrew’s primary focus and the companies journey (08:05)
  • Sidero’s focus on distributed systems (17:48)
  • The challenges of a project that is far down the stack (22:15)
  • Some final thoughts from Andrew (31:35)

Links:

Andrew

  • LinkedIn: https://www.linkedin.com/in/andrewrynhard/
  • Twitter: https://twitter.com/andrewrynhard?lang=en
  • Company:https://www.siderolabs.com

View Details

Today’s episode is a round table with Morgan McClean, Ben Sigelman, and Alolita Sharma, the maintainers of Open Telemetry. Open Telemetry is a high-quality, ubiquitous, and portable telemetry to enable effective observability with a mission to make telemetry as approachable and applicable as possible. Open Telemetry’s values center on compatibility, reliability, resilience, and performance. With these objectives in hand, our maintainers are making waves in open space.

In this episode Morgan, Ben, and Alolita breakdown their individaul involvement in Open Telemetry. They discuss the paths each of them took to end up there, and the varied skillsets the bring to the project. Open Telemetry’s vision and mission is unique in its clarity and precision, and they share some insight as to why. Open Telemetry’s collaboration allows the space for their mission statement to shine through, and as a project before a product, give to the open source community. Check out the conversation for Morgan, Ben and Alolita’s excellent perspectives!

Highlights:

  • Introductions to our three maintainers and their invovlement in Open Telemetry (00:26)
  • Why and how Open Telemetry has such an explicitly clear mission and vision (03:44)
  • Making it clear what Open Telemetry is not (08:20)
  • Thinking as a project, not product (10:45)
  • The pros and cons of working with “frenemies” (17:55)
  • Why Open Telemetry has been successful (27:22)
  • Closing comments on Open Telemetry (32:35)

Links:

Open Telemetry

  • Twitter: @opentelemetry, Ben (@el_bhs), Alolita (@alolita)
  • Company: https://opentelemetry.io

View Details

Tyler Jewell, Managing Director at Dell Technologies Capital, joins me for a deep conversation about the intersection of capital and technology. As a managing director, Tyler harnesses a focus on developer lead companies and the push he makes for those companies when it comes to funding. For Dell Technologies Capital, the focus is on providing the financial support and backing for the market that is developing around the developers themselves.

Tyler breaks down how he honed his focus on backing developers. He refers to the rise of software developers as a “talent class” where he could cultivate investments and partnerships. Tyler shares his parameters for how he categorizes companies and software into four “buckets,” which facilitates the focus he lends to these companies. From identification, to the intersection with capital, check out this conversation for Tyler’s in-depth and exacting definitions.

Highlights:

  • Introduction to Tyler and Dell Technologies Capital (00:00)
  • Defining what it means to be “developer lead.” (02:15)
  • Tyler defines his differences between DevTools and DevPlatforms (4:50)
  • Who is a developer? What is the difference between developer lead companies and the rest? (08:20)
  • Tyler provides insight for those who want to found a developer centric company (14:35)
  • Tyler’s predictions for the coming year, and some advice (21:45

Links:

Tyler

  • LinkedIn
  • Twitter
  • Dell Technologies Capital

View Details

Rob Hirschfeld, CEO of and co-founder of RakN, joins the show to discuss their work in the world of automation. Notably so, automation of data centers using infrastructure code principles to create “infrastructure pipelines.” Rob’s honest and open story provides a great example of how to identify areas that need a product, to developing the product itself.

In this episode, Rob gives us the history of RakN from the earlier inception when he was at Dell, to where they stand today. Rob shares some insight on the challenges of DevOps when it comes to dealing with the various “silos” that organizations have created. He reflects on their transition away from Dell, and how they realized they needed to be table to talk to customers about how they used their products.

Highlights:

  • Introduction to Rob and RakN, a focus on automation, and their origins (00:00)
  • The differences between building and shipping instead of just building on site (04:40)
  • The transition away from Dell and RakN’s early growing pains (08:20)
  • The hardest parts of the technology/commercial balance (14:45)
  • Some critical lessons from the transition (23:25)
  • Reflecting on the early days and lessons (30:25)

Links:

Rob

  • Twitter
  • LinkedIn

View Details

Sam Bhagwat, Co-founder & Chief Strategy Officer at Gatsby, joins me for a conversation about his work. Gatsby is an open source project using React with a focus on building web sites, and not just web apps. As CSO Sam tracks the trends in the modern web development space and helps Gatsby to stay on the innovative edge. From the origin of the open source project in 2015, to the establishment of the company in 2017, Sam and Gatsby’s contributions have only grown exponentially since then.

Sam talks about the history of Gatsby’s rise to prominence and their shift from open source into a proper business. Sam dives into how and why they’ve leaned on investment into the company to help them better address the needs of the web site development ecosystem. From the first service they charged money for, Sam’s take on open source and commercial crossing paths, to Gatsby’s global focus, Sam offers up a lot for consideration!

Highlights:

  • Introduction to Sam and Gatsby (00:00)
  • Moving from purely open source into a business (03:20)
  • Sam’s perspectives on open source and commercial offerings (08:20)
  • Open source as not only a DevTool (12:03)
  • Building websites and brining multiple parties on board (15:16)
  • Sam’s advice to others starting open source projects (21:06)
  • Where to find Sam (25:40)

Links:

Sam

  • LinkedIn
  • Twitter
  • https://www.gatsbyjs.com/

View Details

Avi Press, CEO and Co-Founder of Scarf, joins me for an in depth conversation about Scarf and the work they are doing in transparecny in open source maintainers. Avi’s career and the tools he built lead to a decision to capitalize on his tools. Now Scarf is an extension of his work into a commercial opportunity to change the open source ecosystem.

Avi addresses the general maintaner issues that Scarf wishes to solve. Avi expands on his processes that have landed on a data forward approach and the importance of making that data is a viable capital value. Avi also breaks down the uses of Scarf for maintainers and the suite of tools they are implementing. Importantly so, Avi talks about the ways that the open source space can change to stay innovative and relevant.

Highlights:

  • Introduction to Avi and Scarf’s work (00:00)
  • Shaping his own tools for use (03:50)
  • Scarf’s suite of tools and engaging with the metrics (06:14)
  • Some highlights and information for Scarf users (12:41)
  • Areas where Scarf is building (17:00)
  • How Scarf is working to guard privacy (22:12)
  • Making registry lock in a conversation (27:30)
  • Avi discusses the open source world and some changes that can be made (30:00)

Links:

Avi

  • LinkedIn
  • Twitter: @avi_press
  • Scarf

View Details

Alex Chircop, CEO of Ondat.io, joins me to talk about his company and their recent rebrand to reflect their shift to focus on some fundamental changes in the industry. With a changing persepective that mirrors the changes happening in cloud data and its uses, Alex and the teams at Ondat.io are staying ahead of the curve and implementing some institutional changes.

In this episode Alex goes into the details of their rebranding, and he discusses how they are shifting to answer what Alex calls the “infrastructural dilemma.” With the massive shift to cloud native that developers are making, and the requirements that are demanded by the adoption of these infrastructural demands, Alex and his team are staying in step with the larger community. Alex also discusses the “why” behind their drive to rebrand, and their determination to maneuver the concept of storage as server based into data services that are application centric.

Highlights:

  • Alex’s introdcution and the rebranding of Ondat (00:29)
  • What was happening inside Storage OS that drove them to rebrand (05:00)
  • What it took to follow through a succesful rebranding (09:33)
  • The feedback from the rebranding and branding as a personal choice (14:10)
  • The teams coalescence around the new brand and landing on a name (18:57)
  • Some advice for others who are looking to rebrand (22:33)
  • What is next for Ondat and thier coming services (24:30)

Links:

Alex

  • LinkedIn
  • Twitter
  • www.ondat.io

View Details

Matt Yonkovit, Head of Open Source Strategy at Percona, joins me for a conversation about Percona’s work and their robust history in open source. Percona has been at it for 15 years now and Matt’s work there is both prolific and sets him up to be very well informed about open source strategy at large.

In this episode, Matt discusses what exactly his job means within the context of Percona, and how he covers down to help both the higher echelons of the company, but also the community. Matt provides an excellent bird's eye view of what is going on in the world of open source. His experience highlights many of the challenges that the open source model is currently facing, and can expect to face in the near future.

Highlights:

  • Introduction to Percona and the work Matt does there (00:00)
  • Open source strategy and Matt’s take on its role in an organization (04:20)
  • Choosing where to place priorities and where open source is going (08:00)
  • The ins and outs of pricing models/Percona’s contributions (12:00)
  • Who doesn’t fit the Percona mold? (17:09)
  • Maintaining integrity and staying malleable (21:30)
  • “Shooting [the] sacred cows” of growth (30:45)

Links:

Connect with Matt

  • LinkedIn: https://www.linkedin.com/in/myonk
  • Twitter: https://twitter.com/myonkovit
  • Percona: https://www.percona.com

View Details

Dawn Foster, Director of Open Source Community Strategy at VMware, joins me to chat about open sourcing and the potential risks to consider. With 20+ years of experience in business technology, Dawn lends great insight not only as a leader in the realm of open source, but as a champion for measuring project health.

In this episode, Dawn discusses key risks to consider when open sourcing a project and what startups and small companies should think about as they embrace open source technologies. We also explore trust as a currency of open source, donating to neutral foundations, the CNCF Project Health Measurement Guide, and more.

Highlights:

  • What companies should consider when open sourcing a project. (00:24)
  • Risks associated with open sourcing and the advantage of contributing projects to foundations. (03:59)
  • Dawn explores the interrelation between using and contributing to an open source project. (07:26)
  • A discussion about evaluating and prioritizing a large number of projects - and why smaller companies should be deliberate about the open source technologies they embrace. (11:49)
  • A look at contributor risk with examples of how the risks can vary depending on the project. (17:09)
  • The value of trust in the open source - and Kim’s final thoughts on measuring project health. (22:35)

Links:

Connect with Dawn:

  • LinkedIn: https://www.linkedin.com/in/dawnfoster/
  • Twitter: https://twitter.com/geekygirldawn
  • Fast Wonder Blog: https://fastwonderblog.com/
  • VMware: https://www.vmware.com/

View Details

Joe Bignell, the Kubernetes recruiter at InterQuest Group, joins me for an interesting conversation about the current job market in the Kubernetes space and his role and vision as a talent seeker.

In this episode, Joe and I go into the rabbit hole as we explore the global talent shortage and its impact on the Kubernetes ecosystem. Joe shares invaluable insight and perspective on recruiting for startups, what founders can do to attract talent, and why transparency from all sides (company, recruiter, and candidate) is vital. We also discuss remote work and its increased value and how companies can leverage their Kubernetes talent.

Highlights:

  • What is Joe’s responsibility as the Kubernetes recruiter? (00:19)
  • The current hiring challenges in the Kubernetes talent space. (01:00)
  • Joe’s perspective on recruiting best practices - and his thoughts on companies that seek Kubernetes experts. (06:30)
  • A look at the founder/recruiter relationship, how founders can enhance their recruiting positioning, and the evolution of remote work. (12:31)
  • The pros and cons of recruiting for a startup – and why certain hires can ruin a startup. (19:33)
  • What companies can do to create Kubernetes experts and the role InterQuest plays in closing the talent market gap. (23:06)

Links:

  • Joe’s LinkedIn: https://www.linkedin.com/in/joebignell/?originalSubdomain=uk
  • Joe’s Twitter: https://twitter.com/joe_bignell
  • DevOps for Everyone: https://www.meetup.com/DevOps-For-Everyone/
  • InterQuest Group: https://www.interquestgroup.com/

View Details

Jonathan Ellis, CTO and co-founder of DataStax, has always had a startup mindset. In this episode, Jonathan joins me to discuss his journey and entrepreneurial roadmap thus far.

In our conversation, Jonathan shares how he became involved with the Apache Cassandra project and his transition to founding DataStax. He also shares insight on the importance of hiring a go to market team, why hiring executives proves to be more challenging than engineers, building a company based around an open-source project, and more.

Highlights:

  • Jonathan’s views on his identity as a founder and scratching his coding itch through art. (00:23)
  • A look at Jonathan’s journey from Mozy to the Apache Cassandra project. (05:40)
  • The history of DataStax - and Jonathan explores the benefits of building a company around open source. (11:33)
  • Lessons learned: the importance of implementing a go-to-market team, DataStax Kubernetes adoption, and why hiring executives is a challenge. (15:58)
  • Jonathan’s advice to technical founders - and his perspective and insight on remote work. (27:39)

Links:

Jonathan

LinkedIn: https://www.linkedin.com/in/jbellis/

Twitter: https://twitter.com/spyced

DataSTax: https://www.datastax.com/

View Details

Guy Podjarny, co-founder and President of Snyk, has seen the world of startups through the lens of an employee and as a founder. With two companies under his belt, Guy has excelled as an entrepreneur as Snyk proves to be a leader in developer security.

In this episode of Cloud Native Startup, Guy shares insight on his first startup, Blaze (acquired by Akamai), and his current company, Snyk. He lends perspective on key traits to master when building a successful company and why it’s an exciting time to be an entrepreneur. We also explore why there are major gaps in security, what it will take to fix them, and how Snyk is helping close the gap by decentralizing security.

Highlights:

  • Guy’s interesting evolution from working at startups to finally founding his own. (00:14)
  • Exploring the differences between two worlds: security and DevOps. (05:40)
  • Why Guy sold his first company, Blaze - and lessons learned along the way. (10:27)
  • Guy’s perspective on embracing Snyk’s “failures” - and he reflects on the journey towards identifying their opportunities within the market. (17:27)
  • A discussion about major gaps in security and how Snyke aims to be a solution. (27:39)

Links:

Guy

  • LinkedIn: https://uk.linkedin.com/in/guypo
  • Twitter: https://twitter.com/guypod
  • Podcast: The Secure Developer
  • Snyk: https://snyk.io/

View Details

For Vidya Raman, technology has always been close to her heart. As an investor at Sorenson Ventures, Vidya is guided by this passion and plays an impactful role in helping technical founders build and grow successful businesses. Vidya serves as a leader in early-stage startup investing and thrives on optimizing companies.

In this episode of Cloud Native Startup, Vidya talks about her transition from engineering to venture capital investing, important criteria to consider when evaluating companies, what founders should look for in VCs, her lessons learned, and more.

Highlights:

  • A look at Vidya’s background and the journey that lead to her leap into venture capital investing. (00:11)
  • Exploring Vidya’s role, her passion for partnering deeply with founders, and misconceptions that founders often have about VC. (05:37)
  • What to identify when evaluating companies - and how Viyda measures good fit. (10:41)
  • From the lens of a founder, Vidya shares top criteria founders should consider when seeking VC. (18:58)
  • A discussion about conflicts of interest and managing disagreements between investors and companies. (23:09)
  • Vidya shares her top three lessons learned. (29:23)

Links:

Vidya

  • LinkedIn: https://www.linkedin.com/in/vidya-raman/
  • Twitter: https://twitter.com/veenormous
  • Sorenson Ventures: https://www.sorensoncapital.com/

View Details

Glasnostic is a cutting-edge observability solution that enables DevOps, SRE and security teams to effectively control emerging disruptive behaviors. In this episode of Cloud Native Startup, I chat with Glasnotic’s co-founder and CEO, Tobias Kunze.

As a trailblazer in the world of cloud-native technologies and a two-time startup founder, Tobias brings a wealth of insight. Prior to Glasnostic, Tobias founded Makara, which was later acquired by Red Hat Open Shift. In our conversation, we explore his journey from Makara to Glasnostic and his shift from engineering to entrepreneurship. We also discuss why sales and people management are core skills needed to become an entrepreneur, the importance of actively stepping out of your comfort zone, and the staggering pace of technology.

Highlights:

  • A look at how Tobias shifted his focus from engineering to entrepreneurship. (00:15)
  • Tobias’ perspective on applying lessons learned from his first startup the second time around – and the similar challenges he faced with both companies. (09:00)
  • A discussion about the first dollar and why it is important to step outside of your comfort zone as an entrepreneur. (14:16)
  • Why sales and people management skills are core traits of a successful entrepreneur and startup founder – and why these skills are more difficult for engineers to cultivate. (18:19)
  • Tobias uses air traffic control to illustrate his journey towards founding Glasnostic – and shares insight on current challenges in the technology sector. (26:31)

Links:

Tobias

  • LinkedIn: https://www.linkedin.com/in/tkunze
  • Twitter: https://twitter.com/tkunze
  • Glasnostic: http://glasnostic.com/

View Details

Laurent Gil is by no means a novice when it comes to founding companies. As the Co-Founder and Chief Product Officer at CAST.AI, Laurent marks this as his fourth company. As a repeated entrepreneur, Laurent comes with valuable insight, which he brings to this conversation on Cloud Native Startup.

In this episode, we discuss Laurent’s entrepreneurial journey, which has taken him across the globe and he shares his opinion on why entrepreneurs should hear the word “no.” We also discuss the importance of simplifying product features, the bond he’s built with his co-founders, product-market fit, and more.

Highlights:

  • Laurent shares his entrepreneurial resume and what led to the founding of CAST.AI. (00:12)
  • How a conversation over coffee in France led to building a company in Ukraine – and why rejection should fuel entrepreneurs to keep going. (03:31)
  • Challenging moments Laurent faced while building startups – and how he has been able to work with the same co-founders throughout different companies. (08:58)
  • Laurent explains why simplicity is CAST.AI’s leading principle. (12:49)
  • A discussion about CAST.AI’s pivot to optimizing the one single cloud. (17:16)
  • Defining and measuring product-market fit. (23:21)

Links:

Laurent

  • LinkedIn: https://www.linkedin.com/in/laurentgil
  • Email: laurent@cast.ai.
  • Twitter: https://twitter.com/laurentgil
  • CAST.AI: https://cast.ai/

View Details

As the Co-Founder and Chief Technology Officer at Kong Inc., Marco Palladino takes great pride in the company he has built from the ground up. Marco’s story begins with his move from Italy to San Francisco with no money and a 3-month visa. Today, he and his fellow Co-Founder, Augusto Marietti, have undoubtedly earned that pride.

In this episode of Cloud Native Startup, we explore Marco’s life journey, and he reflects on how his obsession with building Kong led to its innovative success. We also discuss why Kong’s pivot was a risk worth taking, compare building companies around open source, dive into the importance of releasing trust as a technical founder, and more.

Highlights:

  • An overview of Marco’s role as Co-Founder and CTO, how it has evolved, and the origin of Kong. (00:10)
  • Why Kong pivoted from an API marketplace to a technology vendor – and Marco looks back on the wins and challenges in the early days of Kong. (05:56)
  • A discussion about building a company around open source – why Marco embraces mistakes as he reflects on Kong’s global impact. (13:20)
  • Key pieces of advice for technical founders (22:15)
  • Marco talks about his experience as a young entrepreneur and provides insight into why this journey has defined him as a leader. (27:49)

Links:

Marco

  • LinkedIn: https://www.linkedin.com/in/marcopalladino/
  • Twitter: https://twitter.com/thekonginc
  • Kong Inc.: https://konghq.com

View Details

In this episode of Cloud Native Startup, Ihor Dvoretskyi, Developer Advocate at Cloud Native Computing Foundation (CNCF), joins me for a conversation about contributing a project to CNCF. We discuss the benefits of contributing an open source project and Ihor shares insight on the key metrics for success. Ihor also defines each of the three project stages; sandbox, incubating, and graduated, and takes a deep dive into the history of this fascinating open source foundation.

Highlights:

  • A closer look at the history of CNCF and the benefits of contributing a project for both contributors and end-users. (00:18)
  • What it means to donate a project and the responsibilities and continuing services CNCF provides. (08:11)
  • Exploring the different maturity levels of CNCF’S projects – sandbox, incubating and graduated. (12:41)
  • How the CNCF Technical Oversight Committee evaluates projects in consideration for contribution – and Ihor provides insight on the stage that requires the most time and commitment. (18:23)
  • Understanding what a successful open source project looks like and how it is measured. (22:42)

Links:

Ihor

  • LinkedIn: https://www.linkedin.com/in/idvoretskyi/
  • Twitter: https://twitter.com/idvoretskyi
  • CNCF: https:// www.cncf.io/

View Details

In this episode of Cloud Native Startup, Justin Borgman, Chairman and CEO of Starburst Data, takes us through Starburst’s evolution as it speedily makes its mark in the enterprise software space. As a two-time startup founder, Justin illustrates his journey towards founding Starburst, alongside his fellow co-founders, and we explore how they built this unique company. He also shares his advice and lessons learned and we discuss what’s in store for Starburst as the company ventures into its next phase.

Highlights:

  • Exploring how Starburst came into being and what led to the realization that a business would be built around Presto. (00:14)
  • Why Starburst has 12 co-founders and the factors that contribute to the strength of the founding team. (05:53)
  • Starburst’s journey towards raising venture funding. (08:19)
  • The evolution of PrestoSQL and how it became Trino – and Justin shares advice for anyone developing an open source project. (13:49)
  • More on Starburst’s shift towards venture funding and how the founders came to that decision. (21:15)
  • Differences between Justin’s first and second experience as a startup founder and how he applies what he’s learned. (24:53)
  • A look at Starburst’s transition into its third phase. (28:13)
  • Mistakes and lessons learned. (32:02)

Links:

Justin

  • LinkedIn: https://www.linkedin.com/in/justinborgman/
  • Twitter: https://twitter.com/justinborgman
  • Starburst: https://www.starburst.io/

View Details

This week on Cloud Native Startup, I’m joined by Tobi Knaup, CEO & Co-Founder of D2iQ.

In this episode, Tobi provides insight on how D2iQ helps its customers change the world through open source and how it is reflected in their core mission. We also explore what led to the creation of this pioneering technology company and how the cloud native space has changed since.

Highlights:

  • Tobi provides an overview of D2iQ and how it came to be - and he walks through their pivot from Mesos to Kubernetes. (00:11)
  • More on the core mission of D2iQ - and how solving technology challenges at Twitter and Airbnb led to the creation of D2iQ. (04:05)
  • How DQI2 partnered with companies in its early days - and the common mistake that D2IQ and most companies make early on. (13:30)
  • Tobi reflects on his journey from an engineer to founder and CEO - and shares his perspective on the evolution of the cloud native ecosystem. (18:44)
  • Tobi’s advice to the younger version of himself and his fellow co-founders. (28:17)

Links:

Tobi

  • LinkedIn: https://www.linkedin.com/in/tobiasknaup/
  • Twitter: https://twitter.com/superguenter

D2iQ

  • Website: https://d2iq.com/
  • Twitter: https://twitter.com/D2iQ

View Details

HashiCorp’s Co-Founder and CTO, Armon Dadgar, joins me for a conversation on Cloud Native Startup.

In this episode, we focus on open source and how it serves as the core to HashiCorp’s identity. We also explore Armon’s journey towards founding HashiCorp with Mitchell Hashimoto and what the future holds as they both lean into their respective passions. Learn a few keys to cultivating a successful open source community, why some companies don’t rely on this success, his lessons learned, and more.

Highlights:

  • A look at Armon’s role as Co-Founder and Chief Technology Officer and what sparked the decision to build HashiCorp alongside Mitchell Hashimoto. (00:13)
  • Armon shares why HashiCorp began as open source and how that developed into a company. (4:46)
  • How an open source community compares to a paid community - and Armon’s take on bootstraping an open source company. (11:23)
  • Why creating an open source project directly from a closed source project is not the best strategy. (14:43)
  • Keys to building a successful open source community and why this is vital to HashiCorp.(18:14)
  • Armon shares lessons learned during the early days of HashiCorp - and his thoughts on the complexities of being a founder. (23:15)
  • How Mitchell’s decision to step back as an individual contributor allows him to focus on his passion - and more on why they chose to monetize HashiCorp. (29:30)

Links:

Armon

  • Twitter: https://twitter.com/armon
  • GitHub: https://github.com/armon
  • Linkedin: https://www.linkedin.com/in/armon-dadgar/

HashiCorp

  • Website: https://www.hashicorp.com/
  • Twitter: www.twitter.com/hashicorp

View Details

Neil Cresswell, co-founder of Portainer, joins me this week on Cloud Native Startup. In this episode, we talk about the evolution of Neil’s fascinating career, which began at age 17, and how it led to Portainer. We also discuss Portainer’s core ethos of simplicity, open source product, Neil’s predictions on Kubernetes, and more.

Highlights:

  • Neil recounts how he went from self-employed to co-founding Portainer. (00:19)
  • A look at who Portainer was originally built for and the moment he realized it would be a commercial entity. (08:09).
  • Neil dives into the three elements to success in open source product. (11:56)
  • Neil’s advice for someone working on an open source - and a look back at his consulting experience. (14:54)
  • Neil’s shares his lessons learned along his journey - and breaks down some differences in his past and present roles. (20:08)
  • How Neil’s team ensures simplicity - and his Kubernetes predictions. (25:09)
  • Neil shares some of the everyday challenges and advantages of working across different timezones. (29:19)

Links:

Neil

  • Linkedin: www.linkedin.com/in/ncresswell
  • Twitter: twitter.com/neilc_cloud

Portainer

  • Website: Portainer.io
  • Twitter: twitter.com/portainer.io

View Details

Michael Hyatt joins me on this episode of Cloud Native Startup.

Not only is Michael a leading tech investor and philanthropist, but he also ranks as one of Canada’s top entrepreneurs. In this episode, Michael provides a wealth of knowledge as he shares invaluable tips for aspiring, new, and current founders. We also discuss the early stages of founding companies with his brother Richard, the mentality behind hiring your weakness, the phenomenal impact of computing, and much more.

In this episode, we cover:

  • Michael’s take on why starting a company with his brother was a powerful and successful business move. (00:15)
  • The power of computing - and the awful truth about technology. (4:12)
  • The importance of being able to pivot and hire your weakness. (8:18)
  • What Michael looks for when he is investing in a company. (17:27)
  • The impact computing has on new companies now vs 20 years ago. (20:24)
  • Michael reflects on the moment he knew success was on the horizon - and how his inferiority complex played a role. (26:07)
  • The challenges of creating a company built around technology. (28:49)
  • Why you need marque customers. (32:24)
  • Michael’s advice to founders. (35:47)

Links:

Michael Hyatt

  • Linkedin: https://www.linkedin.com/in/michaelhyatt1
  • Twitter: https://twitter.com/mhyattoffice

View Details

J.J. Guy, Co-founder and CEO of Sevco Security, joins me this week on Cloud Native Startup. In this episode, J.J. breaks down Sevco Security and the IT security ecosystem. We also discuss challenges, lessons learned, building a solid team culture and more.

Highlights:

  • Introduction to Sevco and how it fits into the security product ecosystem. (:0016)
  • What creates friction across the entire IT organization and why it is taken for granted. (5:05)
  • J.J. shares ideas that he considered but then ultimately rejected and what he would have done differently. (10:39)
  • The fascinating challenges around enterprise products in security. (17:44)
  • Key lessons learned and how J.J. applies them to Sevco Security. (22:50)
  • The challenges of developing a new product in a new market segment. (27:23)

Links:

J.J.

  • LinkedIn: https://www.linkedin.com/in/jjguy/
  • Twitter: https://twitter.com/jjguy

Sevco Security

  • Website: https://sevcosecurity.com/

View Details

This week on Cloud Native Startup, I am joined by William Morgan, CEO of Buoyant, Inc. In this episode, William talks about his beginnings as a software engineer at Twitter and his transition towards starting and running his own company. We also discuss how rewriting Linkerd enhanced its core value of simplicity, the blessing and curse of open source, his advice to his younger self, and more.

In this episode, we cover:

  • Who is the CEO of Buoyant? William Morgan shares his background. (00:14)
  • The story behind Linkerd and how it came into existence. (1:40)
  • Building a company around Linkerd - and monetizing an open source project. (3:43)
  • William’s thoughts on the open-core model. (7:52)
  • The evolution of Buoyant: Building the long-term future of the business. (9:08)
  • Lessons learned and advice to William’s younger self. (13:48)
  • Deep dive into the process of simplifying Linkerd to reach its core value. (17:34)
  • Istio or Linkerd? William’s take on deciding what is right for you. (23:26)
  • The blessing and curse of open source. (27:03)
  • William reflects on his journey as an engineer to starting and running a company: (32:00)

Links:

William Morgan

  • Linkedin: https://www.linkedin.com/in/wmorgan
  • Twitter: https://twitter.com/wm

Buoyant

  • Website: https://buoyant.io
  • Twitter: https://twitter.com/BuoyantIO

Linkerd

  • Website: https://buoyant.io/linkerd/
  • Linkerd Twitter: https://twitter.com/Linkerd

View Details

David Friend, co-founder and CEO of Wasabi Technologies, Inc. writes the rules of his own success. With 7 companies under his belt, David continues to be an impactful maverick entrepreneur. In this episode, David and I talk about the evolution of his journey, which started off in the music industry, and how it led him to found a cloud data storage company. Join us for more on this week’s episode of Cloud Data Startup.

Highlights:

  • The evolution of David’s companies from the very beginning of his entrepreneurial journey. (00:30)
  • David compares running a synthesizer company to a cloud storage company. (5:42)
  • How careful hiring allows David to focus on his strengths. (10:18)
  • David shares why price and simplicity are the two most important ingredients in selling. (13:50)
  • The influence of data storage and AI and how it has changed the mindset of customers over time. (16:59)
  • David's philosophy on running a business and the joys of conducting the orchestra of his company. (20:46)
  • Advice for first-time founders. (22:18)
  • David’s philosophy on raising money as an entrepreneur. (24:47)

Links:

David

  • LinkedIn: https://www.linkedin.com/in/david-friend-3660832/
  • Twitter: https://twitter.com/wasabi_dave

Wasabi Technologies Inc.

  • Website: https://wasabi.com/
  • Twitter: https://twitter.com/wasabi_cloud

View Details

This week, Swaroop Jagadish, co-founder of Acryl Data, takes us through his journey from quitting his day job as an engineer to founding his first startup company alongside Shirshanka Das. Swaroop also shares his insights on Acryl Data’s business model and the advantages and challenges of building an open source project-based company. Tune in for more on Swaroop and Acryl Data in this episode of Cloud Native Startup.

Highlights

  • How Swaroop created Acryl Data and the story behind the name. (00:42)
  • Lessons Swaroop learned on his journey towards becoming a startup founder. (6:52)
  • Swaroop reflects on the challenges of building an open source project-based company. (10:31)
  • How Acryl Data’s core ethos aligns with LinkedIn (12:49)
  • The unique advantages of Acryl Data’s business model. (14:03)
  • The scariest part of the startup journey. (16:27)
  • Swaroop’s thoughts on approaching the modern data ecosystem. (17:32)
  • Acryl Data’s use-cases. (20:12)
  • Swaroop’s insights on building an open-source company and generally, as a first-time founder (27:32)

Links:

Swaroop:

  • LinkedIn: https://www.linkedin.com/in/swaroopjagadish

Acryl Data

  • Website: https://www.acryldata.io/
  • Twitter: https://twitter.com/acryldata
  • Data Hub Project Website: https://datahubproject.io/
  • Data Hub Slack Community: https://datahubspace.slack.com

View Details

Dhiraj Sharan, CEO and founder of Query.AI, joins me this week on Cloud Native Startup.

With a career that spans over 20 years in cybersecurity, Dhiraj has seen the swift adoption of multi-cloud environments and SaaS apps. In this episode, Dhiraj discusses the importance of evolving your product as the world changes and why you should ask yourself, “How can I be innovative for the next layer?” Dhiraj also gives his younger self advice on the cybersecurity game of chess and much more.

Highlights

  • What led Dhiraj to pursue Query.AI? (0:23)
  • Dhiraj shares ideas that he ultimately didn’t pursue (3:52)
  • The importance of evolving and the layers of innovation. (5:30)
  • The difference between being a founder and being an early employee. (8:52)
  • Advice Dhiraj would give himself if he could go back 20 years. (10:19)
  • Dhiraj’s advice on building a company. (11:52)
  • Security team budgets and creating an ROI calculator. (13:11)
  • Query.AI’s “unique” struggle. (18:47)
  • CSO’s general response to Query.AI. (20:03)
  • Dhiraj reflects on lessons learned and the importance of understanding your customer. (21:48)
  • Dhiraj on the entrepreneur in everyone. (23:32)

Links:

  • Dhiraj:
    • LinkedIn: https://www.linkedin.com/in/dhirajsharan
    • Twitter: https://twitter.com/dhirajsharan
  • Query.AI:
    • Website: https://query.ai
    • Twitter: https://www.twitter.com/query_ai

View Details

Anurag Goel, Founder and CEO of Render, joins me on this episode of Cloud Native Startup. Learn about his beginnings at Stripe as employee #8, the birth of Render and how it solved a gap in the market, and his lessons learned while transitioning roles. We also discuss how open source fits into business strategy and much more.

In this episode, we cover:

  • Anurag and how he founded Render (00:27)
  • Lessons learned while transitioning professional focus from engineering at Stripe to strategic business (05:42)
  • Anurag’s advice for people who want to transition into larger business roles (11:11)
  • The evolution of Render and its market and use case (13:57)
  • The important role of open source in the developer tool ecosystem (19:46)
  • Why Anurag chose to make Render open source (22:11)
  • Strategies for making open source part of your platform (22:51)
  • How open source fits into business strategy (23:55)

Links:

  • Anurag: www.twitter.com/anuraggoel
  • Render
    • Website: https://render.com/
    • Twitter: https://www.twitter.com/render

View Details

This week on Cloud Native Startup, I talked with Wei Dang, founder of cloud native security company StackRox which was acquired by RedHat in 2020.

Highlights:

How Wei met his co-founder and how the two of them saw the need for new types of security tools.

Why talking to people throughout the Kubernetes ecosystem led to a series of a realizations that security in a cloud native world was going to me an increasingly important part of the conversations as more people adopted Kubernetes.

Where the name StackRox came from.

How even understanding if there was a market for a container security product. The moments wondering ‘are we building the right product’ was the scary.

Why it’s important to focus at the beginning.

How StackRox evolved from container security to Kubernetes security as the broader conversation shifted and the industry consolidated around Kubernetes.

The moment Wei felt like there was product-market fit for StackRox.

How Wei would define Kubernetes Security.

The ways in which starting and growing a company forced Wei to learn new skills and gain knowledge.

Why community is so important for companies in the Kubernetes ecosystem.

How things have changed — and how they haven’t — since becoming part of Red Hat.

Links

https://twitter.com/weiliendang

https://www.linkedin.com/in/weiliendang/

View Details

The Business of Cloud Native is now Cloud Native Startup. Going forward, I'm moving away from talking to end users and focusing instead on what it takes to build a startup in the cloud native ecosystem. I'll be talking with startup founders, startup advisors and others in the ecosystem about making the transition from software engineer to startup founder, the stories behind companies we read about in the tech press and how to increase your odds of success in the cloud native startup world. 

View Details

This week on The Business of Cloud Native, I spoke with Abby Kearns, CTO at Puppet, about the changing role of technology in the enterprise and how that changes things for software companies like Puppet.

Highlights

Whatever the exact definition of cloud native, ultimately it is a way to improve scaling and resilience.

The increasing importance of software for enterprises, because customers simply expect to use software to connect with companies of all sizes.

Why cloud native is tied to digital transformation — because you can’t do one without the other.

Picking tech and deploying it — that’s the easy part. But the people part is hard. Changing organizational structures is hard, but the companies that are succeeding in the new digital environment have to do the hard work.

Why enterprises that are successfully using software to build a better relationship with your customers have top company leadership, from the CEO on down, investing in the transformation and incorporating technology into the company’s vision.

How the changing role of technology in the enterprise has changed things for technology companies like Puppet.

How technology decisions have now become board-level decisions, rather than decisions that made in a basement among technologists.

How the changing landscape has forced Puppet to change its go to market strategy, positioning and messaging.

The speed of change in the technology space seems to be accelerating and can lead to a lot of uncertainty.

Why open source can be extremely rewarding, but requires companies to give up a huge amount of control that can be unnerving for enterprises.

What is the role of a CTO at a technology company — is it tactical? Is it visionary?

Links

Abby Kearns on LinkedIn

Abby Kearns on Twitter

Puppet

View Details

This week on the Business of Cloud Native, I talked with Kelsey Hightower, principle engineer at Google and Kubernetes expert. We talked about how technology like Kubernetes helps engineers focus more on solving business problems instead of constantly solving the same low-level problems.

Highlights:

How the evolution of technology — and the evolution of customers’ expectation — have made cloud native practices table stakes.

The more established a company is, the more layers it likely has in its technology stacks. Unless a company is under 10 years old, it probably isn’t 100% cloud native because it’s rarely practical to throw away everything they’ve been doing in the past.

Why it’s so challenging for companies to “disrupt themselves” by adopting cloud native technology unless they have serious motivation.

How successful cloud native journeys involve both grand strategic visions and boring tactical plans that can actually be implemented.

Why companies need to take into account the entire ‘infrastructure’ needed to adopt cloud native. How do you collect the data you need to reach your goals? Do you have the human resources, both technical and non-technical, to achieve their goals?

How cloud native transitions can quickly become an ‘onion’ problem where there is always another layer that companies need to solve.

How to convince practitioners who are trying to build customer tools internally that they should use Kubernetes or other open source projects.

The myth of ‘tech displacing people.’ Usually evolution of technology leads to more jobs for software engineers, not less.

How Kubernetes helps engineers focus on the business — and that is a good thing.

The difference between 20 years of experience and 20 years worth of 1-year experience.

Why Kubernetes is a platform for building other platforms.

Why Kubernetes and cloud native are not a magic bullet that will completely transform your business.

Links:

Kelsey Hightower on Twitter

View Details

This week on The Business of Cloud Native, I spoke with Mark Thiele of Edgevana about the definition of cloud and cloud native, where the line is between cloud and edge and situations where edge is the best option.

Highlights:

A discussion of situations where edge is the most appropriate option and how edge can help solve problems related to latency, data transfer costs, data sovereignty and network access.

Why the right architecture depends on your business needs — there is no one size fits all way to design an edge solution.

The relationship between data centers, public clouds and edge devices, including how co-location facilities fit into the equation.

How IT is shifting from being a cost center in the business to being a profit center, and how this re-framing of IT is the root of fundamental change.

What types of companies can successfully manage a data center and what types of companies should accept that they don’t have the skills and just go to the public cloud.

Why success in the digital transformation is really a question of prioritization.

Do you enjoy the podcast? Help others find it by leaving a review on Apple Podcasts and sharing on social media.

Links

Edgevana

Mark on LinkedIn

Mark on Twitter

View Details

This week on The Business of Cloud Native, I spoke with Jana Boruta, Director of Global Events at HashiCorp, about building community — for startups, for big companies and for open source projects with no budget. We also touched on digital-first events and how they differ from in person events.

Highlights:

What is community and why does it matter for companies building commercial tools for developers.

Why community building is for everyone — from Nike to volunteer organizations to enterprise software.

Is it ever too early for community building? How to determine if you’re building community for the right reasons.

Why community building is a long game — not something that will show immediate ROI after a single event.

Community building starts with creating a community blueprint. While some tactics are common, but every company’s community is going to be different and has to be authentic to the company.

Why product feedback can be an important first step for building a community.

Why starting with community building too early, when the product is too buggy, can be counter productive.

Why you should avoid thinking of your community as a demand gen program.

How create digital events that provide value for the attendees, even if they deliver it in a very different way than an in-person event.

Links:

Jana’s website

Jana on LinkedIn

HashiConf

EpicConf

Digital-First Events, Jana’s book about digital events

View Details

This week on The Business of Cloud Native, I spoke with James Campbell, CEO of Cado Security, about his background in the security world and why he felt like there needed to be a better way to manage security forensics in a cloud native environment.

Highlights:

Why it’s important to get better information about security incidents — or potential security incidents — to make better decisions.

Why security has to be relatively easy because otherwise people will ignore it — at their peril.

How cloud native features like auto-scaling are great for compute but make security, especially security forensics, more complex.

Without enough data collected in real time, companies can end up unable to know whether or not an anomaly actually caused data loss, which data was impacted and what the root cause of incident was.

How some of the most sophisticated attackers operate and how they can cause havoc even if the impacted container has spun down.

The triggers that led Campbell and his co-founder to start Cado Security.

Why having better information is critical to responding effectively to breaches, large and small.

Links:

James Campbell on LinkedIn

James Campbell on Twitter

Cado Security

View Details

What is edge? What is cloud? What is the difference and what are the different requirements for each use case? I talked to Samy Fodil, CEO and founder of Taubyte, about edge computing, industrial IoT and how the edge requires a dramatically different approach from the cloud.

Highlights:

Why edge is a truly distributed system, unlike the cloud.

What inspired Fodil to start Taubyte and why he thinks the edge will be the default computing platform in the years to come.

What overlap there is in skill sets between on-prem development, cloud development and developing edge-native applications.

Why edge-to-edge communications can be so hard for developers to understand.

What barriers prevent companies from taking advantage of the edge and why it’s worth it for those businesses in spite of the complexity.

Links:

Samy Fodil on LinkedIn

Taubyte

View Details

This week on The Business of Cloud Native I spoke with Travis Nielsen and Sebastien Han about how Rook has evolved over the years, beginning as a way to build a cloud native storage platform but before anyone was talking about cloud native or Kubernetes.

Links:

Travis on LinkedIn

Sebastien on LinkedIn

The Rook project page

View Details

This week on The Business of Cloud Native, I talked with former Gartner Analyst, current VP of Solutions at Securonix Augusto Barros. We talked about how Securonix’s positioning has evolved over the years as it has move between market categories as well as how to evaluate and test cloud native security solutions.

Links:

Augusto on Twitter

Augusto on LinkedIn

Securonix

View Details

In this episode with Evan Reiser, CEO of Abnormal Security, we explored how positioning and an understanding of who your ideal customers are and what their needs are can influence technology choices, including which cloud provider you build on.

Highlights

How a seemingly pure technology choice like cloud provider can have serious implications for customer experience.

The difference between having a board-level discussion about cloud providers is different from gathering the engineering team to talk about cloud infrastructure

Why being integrated in the Microsoft ecosystem was a strategic business decision and how technology decisions in general can be high-level business decisions

Why technology teams should think more about what the customers need and want instead of just choosing the best tool from a technical perspective

Evan’s hesitations about making the transition to Azure and why they did it anyway

Why they chose to re-architect at the time they did

Even though the move to Azure was made to improve customer experience, customers don’t necessarily have a different experience since the move

Why founders should keep in mind that startups rarely fail because their technology doesn’t work, but because they don’t meet the needs of their customers

Links:

Evan on LinkedIn

Abnormal Security

View Details

This week on The Business of Cloud Native I talked with Yong Tang, one of the maintainers of CoreDNS, about how the project started, how it’s evolved over the years and how the team decided to integrate it with Kubernetes.

Links:

CoreDNS

Yong Tang on LinkedIn

View Details

This week on The Business of Cloud Native, I talked with Brian Gracely about using Kubernetes for edge workloads as well as the difference between “cloud” and “edge.”

Highlights

Is edge part of the cloud, is cloud a part of edge or are they completely separate but slightly related environments?

What makes something a data center vs what makes something an edge device?

How enterprises think of edge vs how Telcos think about edge.

How the edge has gone from being a cost center to a competitive advantage for enterprises.

Why Telcos have always thought of edge as a market opportunity.

Why Kubernetes can help standardize environments and make it easier to deploy software to the edge, but there are still challenges to overcome.

Why edge deployments require re-thinking many basic environmental factors like bandwidth and compute capacity.

Why consistency at the edge is so important.

Why you can’t ignore the physical conditions that make edge environments unique.

Links

Brian on Twitter

The Cloudcast

RedHat OpenShift

View Details

In today’s episode of The Business of Cloud Native, I talked with Julien Pivotto and Richard Hartmann, two of the maintainers of Prometheus, about how the project started, how it’s evolved over the years (and how it’s stayed the same) as well as some novel ways Prometheus is used in the real world.

Highlights

Why both Julien and Richard got started with Prometheus

Some surprising ways that Prometheus is used to monitor things beyond the software engineering world

How Prometheus has evolved in technology and usage over the years

How Kubernetes and its relationship with Prometheus has changed the project

What assumptions ‘cloud native’ creates for potential Prometheus users

Links

Julien on Twitter

Richard on Twitter

View Details

In this episode of The Business of Cloud Native, Chris Holmes talks about bootstrapping Decipher Technology Studies and their core product, intelligent service mesh Greymatter.io. He also talks about why it's so important for brownfield and greenfield apps to talk to one another and the many similarities between public sector and private sector organizations.

Highlights:

How Greymatter combines business intelligence and security controls.

The difference between working with public sector customers and private sector / enterprise customers — and why there are more similarities than differences.

How segmentation is sometimes necessary for any highly security-conscious organization, including both government organizations and financial services companies in the private sector.

Why we need to respect legacy applications — because they tend to be the mission-critical applications that drive revenue.

Why connecting brownfield and greenfield applications is critical, because not all ‘legacy’ apps will ever be moved to the cloud.

What ‘returns’ a company is looking for when evaluating ROI on cloud migrations.

What we mean when we talk about an “ROI” on security tools.

Why Kubernetes’ terrible networking is part of why Chris could see that service meshes would be necessary even back in 2015.

Links:

Chris on LinkedIn

greymatter.io

View Details

This week on The Business of Cloud Native I spoke with Chris Psaltis, CEO and co-founder of mist.io. We spoke about why multicloud is necessary (and scenarios where multicloud is not necessary), where multicloud is headed in the future and the journey Chris and his co-founders have been on with Mist.

Highlights

The difference between using multicloud for legal / regulatory reasons or because of the company’s history and using multicloud strategically to improve developer velocity or improve customer experience.

The complexity involved with pursuing multicloud and why many organizations are better off in just one cloud.

Why being cloud agnostic from day one is not a good strategy in the vast majority of cases.

Why no one seems to be able to correctly estimate how difficult it is to build a multicloud platform.

Why ‘silos’ are the competitive alternative to a unified platform for companies

How Mist went from working primarily with smaller teams before figuring out that they provided more value for large teams because the pain from multicloud management increases exponentially as the number of engineers, applications and environments increases.

When the founding team decided to stop being consultants and start an open source technology startup.

Links

Mist

Chris on LinkedIn

Chris on Twitter

View Details

This week on The Business of Cloud Native, I talked to Tzury Bar Yochay, founder and CTO of Reblaze, about building a cloud native security company before twelve thousand people were going to KubeCon.

Highlights:

Why your security measures have to keep up with hackers’ sophistication.

The moment when Tzury decided to go from being a contractor for the defense industry to founding a company.

Why the default path for startups is failure.

Why open source is key to securing your cloud environment.

How selling a security product to developers has evolved over time.

Why lazy developers are good developers.

Why selling software to developers is different from selling software to other types of professionals.

Why he thinks the most brilliant developers tend to gravitate towards open source.

Why security based on obscurity is a terrible, perhaps even evil, strategy.

Links:

Reblaze

Tzury on LinkedIn

@tzury on Twitter

Tzury on GitHub

Curiefense

View Details

This week, I talked with Karthik Ranganathan about the challenges going from employee of a large company to startup frounder and why he founded Yugabyte because he wanted a database that both was transactional and still could be highly available.

Highlights:

Why the ability to scale is important for any cloud native application, including for a cloud native database.

Why Yugabyte is still open source and why being open source is important to the company.

Why enterprises wanted an open source database to house their mission-critical data.

Why the company went from an open core model to open source / managed service model.

Why end customers care about open source.

Why early-stage, small companies have trouble establishing trust and how being open source helps build trust.

Why building around open source helps nudge customers to ‘buy’ instead of built it themselves.

Why finding the right position and the right message is a major challenge at the beginning of the company.

Links:

Karthik on LinkedIn

Karthik on Twitter

Yugabyte

Yugabyte Slack

View Details

This week, I talked to Natalie Ledbetter, Head of People and Platform at Boldstart Ventures. We talked about how startups can approach team and culture building, including:

  • How to prioritize your hires
  • Common mistakes founders make when building a team
  • Why you should always avoid brilliant jerks, even if they are very brilliant
  • How to divide responsibilities between founders
  • Anticipating growth and setting your team up so that it can scale as easily as possible
  • The difference in skills sets between 'startup people' and employees you would want to hire later on

Links: 

Natalie on LinkedIn

Natalie on Twitter

Boldstart Ventures

View Details

In this episode of The Business of Cloud Native, we talk about the hard business goals behind words like "freedom" as well as what it's like to go from engineer to CEO. My guest, Sirish Raghuram, is the CEO and co-founder of Platform9.

Links: 
https://platform9.com/
https://www.linkedin.com/in/sirishraghuram/

View Details

In this episode of The Business of Cloud Native, Shawn Lankton talks about how Microsoft 365 and related applications fit into an organization's move to the cloud and why organizations need to pay attention to security for all their SaaS applications. 

Links:
https://www.coreview.com
Shawn on Twitter

View Details

Thomas Markey talks about how to remove some of the barriers to running an open source project with free hosting services.

Links:

https://fosshost.org

Thomas on LinkedIn
FOSSHOST on Twitter
Discord

View Details

Ranjan Parthasarathy talks about why separating compute and storage makes it easier to operate at hyper scale and why he decided to found Logiq.ai to make it easier for companies to do so.

Links:

LinkedIn
Logiq.ai

View Details

Have you thought about how your OSS license could impact your ability to grow your community and monetize your OSS in the future? Attorney McCoy Smith talks about what to be aware of at the beginning to avoid messy legal issues down the road.

Links:
McCoy on LinkedIn
Lex Pan Law
Opsequ.io

View Details

As an analyst at CCS insight, Bola Rotibi gets a birds-eye view of trends in how industries use software to advance their business goals. We had an incredible conversation about how companies use cloud native technology to meet business goals and how vendors in the cloud native space should pay more attention to the needs of specific industries. 

Links: 
https://www.ccsinsight.com/blog/author/bolarotibi/
https://www.linkedin.com/in/bolarotibi/
https://twitter.com/bolarotibi

View Details

SaaS companies that handle customers' sensitive data need to worry about how they manage data locality. In this podcast, Canopy CTO and founder talks about how their business would not be possible without the flexibility and ability to easily spin up resources in regions across the world that a cloud native architecture offers. 

Links: 

https://www.canopyco.io
https://www.linkedin.com/in/oransears

View Details

Here's what I covered in this episode: 

  • What positioning and market segmentation is and is not
  • The specific positioning challenges facing companies in the cloud native ecosystem
  • Why it's important to identify and talk about the types of application your product benefits the most

Thanks for listening, and happy new year!

View Details

Jim Bugwadia, CEO and co-founder of Nirmata, talks about what has changed (and what has stayed the same) since the company started in 2013. 

Links:

https://www.linkedin.com/in/jimbugwadia/
https://nirmata.com
https://kyverno.io

View Details

In episode 29 of The Business of Cloud Native, I talked to Krishnan Subramanian of Rishidot Research about trends he sees in how end users use cloud native technologies and how startups in the space can meet end users where they are.

Links:

https://rishidot.com

https://www.linkedin.com/in/krishnansubramanian/

https://twitter.com/krishnan

View Details

This conversation covers:

  • Idit’s role at Solo.io, and what she typically does on a daily basis. Idit also talks about how her job duties have changed over the last two years, and the impact that COVID-19 had on the company.
  • The common business reasons why customers come to Solo.io — and where they typically are in terms of cloud-native maturity.
  • Some things that Idit has learned about customers over the last two years. In addition, Idit talks about her own journey at Solo.io and what she’s had to learn along the way.
  • How Idit’s customers typically benefit from using distributed systems — and some of the top misconceptions that they tend to have about using them.
  • Idit’s thoughts on the market for cloud-native technologies.

Links

  • Solo.io
  • Follow Idit on Twitter
  • Slack

TranscriptEmily: Hi everyone. I’m Emily Omier, your host, and my day job is helping companies position themselves in the cloud-native ecosystem so that their product’s value is obvious to end-users. I started this podcast because organizations embark on the cloud naive journey for business reasons, but in general, the industry doesn’t talk about them. Instead, we talk a lot about technical reasons. I’m hoping that with this podcast, we focus more on the business goals and business motivations that lead organizations to adopt cloud-native and Kubernetes. I hope you’ll join me.

Emily: Welcome to the Business of Cloud Native. I'm Emily Omier, your host, and today I'm chatting with Idit Levine of Solo.io. Idit, I want to start out, first of all, by thanking you for joining me.

Idit: Oh, thanks so much for having me.

Emily: And then, second of all, I wanted you to just start off by introducing yourself: what you do, what your company does, and also a little bit about how that translates into what you do every day, like, what activities you spend your day doing.

Idit: Oh, for sure. Okay, so as you said, my name is Idit Levine. And I’m, right now, the founder and the CEO of Solo.io. I started Solo two years ago, and when I started it, my focus was try to solve our [00:01:24 unintelligible] application networking problem that we know that will come up.

So, what does it mean? As you guys all know, there was a huge shift in the market between monolithic to microservices and, kind of like, moving from technology of monolithic to microservice stack mean that now we also moved to a distributed application. And it was clear to me that now everything is basically will go on the wire; any communication, small communication, between those two microservices basically will have to go to the network. And I thought that would become a big problem because stuff that we didn't need to take care of when everything was the same binary, now we need to actually figure out how to solve. And basically, I was really passionate, thought that that will be a huge problem in the ecosystem and I was very passionate to actually try to solve that. So, the idea was, how to connect, right? How to connect the application, how to connect everything related to your, eventually, application to the user.

Emily: And then tell me a little bit, what do you do every day? When you start, what does an average day actually consist of?

Idit: Oh, wow. So, it's really interesting, that I think it's a huge difference between now and what I was doing a year ago. Right now, basically, it's pretty simple. Corona came by and it was influence a lot of companies. I was assume that it will influence also my company, and therefore I basically freeze hiring, freeze everything, and try to do the best I can with the resources that we had.

What happened is that actually, not only that we didn't was influenced, we actually over doubled our revenue every quarter. That's basically forced me to immediately grow the team to be able to actually serve all those customers. Right now, basically, the main thing that I'm focusing on is—besides the technology, of course, in the strategic of the company—is basically on growing the team. So, it's hiring, it's interviewing, it's looking for the right people, it's building. You know, basically try to grow the team as much as I can in order to basically, yeah, serve well, the customer that are asking for us to—you know, for our products. That's a lot of my focus this day.

Emily: And what do you find are the business reasons? What's the business problems that cause somebody to come to you?

Idit: So, as I said, once people basically is moving from monolithic to microservices, there is a lot of simple stuff that before that just natively happened inside of the organization; right now, it's a little bit more complex. So, first of all, they needed to find something to run it on, and this is what Kubernetes so great in this ecosystem is the ability to install, upgrade, and basically orchestrate their microservices. But then, as I said, simple stuff that before that people were baking into the microservices created a lot of issues, like small stuff, like how do two microservices communicate with each other? How do you make sure that they're doing it safely right now? Because as right now, it's all on the wire, so potentially, there's always a third party that could, you know, join the party.

So, you really need to be safe and make sure that there is a very secure line between those microservices. And then the last thing is that because there is so many because the idea of microservices was to allow you to scale, the question is how do where the request is actually routed? So, in the [00:04:52 unintelligible], request is coming, and there is a lot of replication of the same microservices, and you have no idea basically where it's coming and where it's landing. And then it will go to the next level of the microservices, and again, not know which instance of it is basically being hit.

So, now the question is, how do you get visibility to something like that? How do you know what's going on in your cluster? How do what to look for the logs when now it's distributed all over the place. So, that's a lot of problem that the organization basically started to have. As well as with this—if—before that, there was a technology called [00:05:26 api-get] that was relatively popular, but people somehow—it wasn't a must.

Right now, when microservices was adopted specific in environment like Kubernetes, when everything is very cloud-wise, you know, stuff is coming up and coming down, you really wanted to make sure that you have a place that you can actually control the policy, control the [00:05:50 unintelligible], the [00:05:51 unintelligible]. And that's basically where API can help. So, that is basically—how do you manage all this networking, basically, of all these systems and applications, as an edge gateway? It's something that going inside your cluster, as well as what's going on inside the cluster after it. And that's basically, yeah, the main problem that you're solving.

So, every traffic to your infrastructure, node to start, we're basically taking care of exactly of everything that you then have traffic between what called East and West, inside your cluster. And that's basically the stuff that we are targeting. So, customers that calling us is mainly usually aware of more like, API gateway and service mesh, and they are basically looking for the one that is the most innovative one, the one that can solve them the problem in a very Kubernetes-native way. And this is where we actually very attractive.

Emily: And how do you usually position yourself on the cloud-native maturity journey? When somebody comes to you, how far are they, usually, on that maturity journey.

Idit: So, that's very interesting. When I started the company, it was clear to me that service mesh would be a very big thing that everybody will adopt, but I also knew that it will take quite a while until the market will get there. And I also knew that every company that will come with me will be, as you said, in a different level of maturity. So, what's always very focusing on, is basically, is the journey. So, you have customer that's just have monolithic application, don't even using something like Kubernetes yet, and just interesting of basically moving. So, we will have those guys.

And then once they actually adopted, they will need some API gateway that is natively running on Kubernetes and will help them with. And then when they will decided that they have more and more microservices, they probably want to adopt something like service mesh. So, we have that, the second pillar of the company that basically focusing on service mesh, and then once they do this, we even going the next level, and basically help them extend those measures. And that’s, you know what, basically imagine that because of this technology, it's allowing us to basically attack more use cases. So, that's the next level that we are in.

So, the idea is that we're getting customers from all over the [laugh] the spectrum: people that just starting, people that wanted to do already adopt and looking for the next thing, people that adopted and looking for more innovation stuff. So, it's really, kind of like, a journey that you're taking with us, and it doesn't matter where you are in this journey, we usually have a way to help you getting even more innovative.

Emily: And then what do you find about when somebody is going to be happy with say, a pure open source solution to something versus when they need an enterprise version? Like, when they need to pay for something? What are the different characteristics of someone that's going to be happy with pure open source versus not?

Idit: So, specifically now [00:08:47 unintelligible] is very depends in— what is the feature that you're looking for? [00:08:53 unintelligible], we have a lot of people that are using open source are not paying us a dime, and this is totally fine, and we excited about having them as a user. The way we, kind of like, differentiate our model is the thing that we are putting on top of service mesh, it’s what people usually are paying for us. So, stuff that related to security. So, if you really care about security, you want something like web application firewall, you want something like data loss prevention, that's usually when you will basically want to come to us and basically ask for help.

The other thing that is very important is support. So, we have a very, very active open source community. And we are helping as much as we can, and I think we are very, as I said, not only that we are helping right now, our community is helping, which is a beautiful thing to see. But if you are a huge organization, like Vonage, or at a company like ADP is that basically running in production, most likely wanting 24/7 immediately getting support. And that's basically what will be on offer as [00:09:57 unintelligible]. So, it's mainly very related to this.

In terms of the people, I will say that there is a lot of company that just come and gets a set and taking the open source usually is people—as I said, even people that are less familiar with that, we are doing a lot on the [00:10:11 unintelligible], so we can help them to get where they want. I would say that the only thing is basically mainly is this support and some security feature that we're putting on top of it that is not available to open source.

Emily: What have you learned about your customers over the past two years? Do you think there was anything that—like any misconceptions that you have that you’ve discovered that were not correct?

Idit: Yeah, my gosh. Yeah. See, I mean, first of all, [00:10:40 unintelligible] is a little bit interesting. We actually, all the work that we're doing, customers coming to us, it's all inbound leads. So, we are not going to any customer, they are coming to us.

I think that the majority of the thing that I learned that is extremely important is that if customer is committing, and actually starting a POC, you learn so much about these environment, what does it mean for him to run? What does it mean for him to upgrade? Because in the [00:11:06 unintelligible], when you using open source specifically, when there is such a different variety of customers, I don't think you getting enough exposure to the people that running in more big organization because usually those people is not going to use open source; they will look some provider to help them adopt that. And therefore, you're not exposed to their environment, you don't know what their limitation, you don't understand how they—you know, the politics in their organization, you don't understand how the security division is working together with architecture, and so on. And I think just learning this is teaching you a lot because it's explain you what is the limitation in your product that you just, to be honest, didn't even envision that exist.

So, why people is not just taking the new version of my product? This is the best, which is put a lot of a work on this, and why people is not upgrading. Then you learn stuff, like for instance, what does it mean to upgrade in organization, in big organization like this? And you discover that it's a really, really big deal, and it's very hard. And their probably doing it over six months to a year testing before they actually putting into production.

Just to upgrade, it's not something that is extremely simple. And I think all this stuff, if you're looking at the open source community, I think that's what they're missing, all this exposure for different environment, different organization, how does it work? What is the limitation? And then of course, if they don't know, it's very hard to try to solve it. So, I think that that's what make up for the exposure that we have with the customer.

That's what make the products better. And I think that in open source community, it's not always the case. Specifically, not when it's young, relatively. I'm not talking somthing like Kubernetes, that has Red Hat and all those guys involved, and therefore they're basically bringing the information there. I'm talking about open source that basically people start using, usually, they're missing all this input from the customer, which is extremely important.

Emily: And how have your job duties changed over the past two years? One of the things you mentioned was like, “Oh, gosh, what I'm doing now is really different from what I was doing a year ago.” So, what were you doing a year ago? What were you doing two years ago?

Idit: Yeah. So, I think that in the beginning, when we started the company, the focus was more on the product in the open source. My job was basically understand why do I think that—basically to find a market fit. That was basically my focus. I was doing a lot of product stuff, I was trying to understand—of course, we grow the team, and but not in that based because we didn't have customers back then, of course, so was less interesting to me.

Marketing was totally different, with the idea was to create an awareness versus right now we already have awareness, so the marketing is changing more to target leads. So, a lot of stuff in the company changed, totally. How do you manage, right? When we had before, and we were very young, and we have—I don't know—10 engineers, it's totally a different skill to manage when you are 50 engineers. So, a lot of stuff change, and of course, every phase of the company is different.

I think that, as I said, what as the big kind of like, push back and forward was once we started to acquire a lot of customers, and the knowledge base that they gave us; that was extremely useful. That, I think, when you need to start suddenly thinking about stuff that you didn't in the beginning, how do you support that? How do you make sure that they have a 24/7 support to it? How do you help them to actually go to the production? The skill are different.

It's not enough—not necessarily, the engineers are the one who will need to do this. I mean, someone can write the best code in the world, but maybe it's not the best solution architect that exists. So, just attacking different groups. Suddenly, instead of just hiring an engineer, now I needed to hire a solution architect, the field engineers. So, the focus of the company doesn't change; you want people to run into production, want to make sure that we are—you know, it's not only about creating a community, it's not only about open source, it's more about how we’re actually making our product the most usable for the real use cases in the world. And not just people, maybe just for fun trying it in the open source.

Emily: What do you think are the top three things that you've had to learn?

Idit: Yeah, a few things. So, first of all, I’d start with the fact that before I started Solo, I worked in a very big company, EMC. When I started Solo I—one of the perception that was changed for me is that, when you're working in a small company, it's way harder for you to get a lot of stuff that you assume that would be—that was extremely easy when you work in a big company. Like when I was in EMC in the beginning, everybody wanted to interview me. Everything that I did, any announcement that I make, created such a huge noise.

Versus then, when I moved to—started Solo and we were a very small company, and no one know us, everything that I tried to do was [laugh] way harder to get, right? Attention to the market, attention from the industry, attention for reporters, everything was it will be harder to do. So, that's one thing. Totally different. And that's true for everything.

I mean, if Google will go and put some open source project out there, no matter how good it is, it will get a lot of eyes on it and people will be extremely excited. I can put an amazing project to, and the same equivalent project in the open source, most likely, I'm not going to get the same coverage. So, that's number one that basically are saying, “Huh, that's different.”

The second thing that I learned is that I think that a lot of us, not only me, has this perception of, we're going to put a very good project out there, and that's it; that's all we need: good technology out there, people will see it, they will understand how amazing it is, and we will become the best—the next—I don’t know—HashiCorp. That's really not how it works. Also, it's not enough to put the project, and then hope that the community will come and build it with you. As I said, that's maybe true when you are Google; that's really not true when you are Solo.

When you putting a project, doesn't matter how good it is, there’s one of two things that will happen: either people will come, and then just expect you to continue building it, but it will be on you to build it, which is totally fine, and that that's what I think it should be; Or even worse, other people will come and copy the project from you. And you saw that in the open source. So, I think that a lot of the perception about open source, I was pretty naive about how these things work. And I think that was something that I definitely needed to learn.

And I think that the last thing, and as I said, is the most important is exactly understand that in the [00:17:37 unintelligible], I need to get a real customers; people that has skin in the game. They pay you money, and because they pay you money, they will put a team, and they will work, and they will stress it, and they will make sure that it will run perfectly, and they will make the product better by giving you an input. And I think again, in the beginning it, I was relatively naive when I came, and said, “Oh, we’ll put a project. It will be cool. We are smart. We can figure out what to do.”

But actually, it's not working like this because no matter how smart you, if you don't have the data, you can’t make the right decision. I think that that's the third thing that, basically, it was a huge change in the company once we start to get a real customer running it in production. I think it's totally changed the company, and the product, and the view. Everything was extremely different.

Emily: Yeah, that makes total sense. What would you say is the top symptom that a company would experience, that you can help with?

Idit: So, as I said, for us, it's a little bit different. They are coming to us. And then, as I said, they are very different in the level of knowledge. There is customer that coming, and most of the work they're doing by themselves, they already fluent in Kubernetes, they understand even the concept of service mesh. Sometimes they even going in, already installed the open source by themself and start working on this, and then just coming to get fine-tuning of it.

When there's customer that comes in that doesn't know anything. They just know that they have—they wanted to change. They just know that Kubernetes is probably something that they will be interested in, and therefore they will need other stuff. But that's basically where it's done; they don't know much more. And for those people, a lot of the support that we are giving them is not even related to our product, it's basically, just teach them how to use something like Kubernetes.

So, I think that this is a very, very different. In the [00:19:23 unintelligible], as I said, as a customer—as a company what we own is basically application network layer. Everything, as I said, North to South, East to West, and I think that this is something that we extremely solid at this. We’re running Envoy in production for the last two years. We’re running a huge organization in a big scale. We writing our own filter to Envoy, we basically—we doing a lot of upstream code, so you know, it’s putting us as really understanding that piece of the low-level application networking. So, I think that's really helpful.

Emily: Why do you find that your customers summers want to use distributed systems in the first place?

Idit: I’d say it's good question. I think that mainly the reason is usually for scale, and I think that another reason—and I will not say that this is a distributed system—but the ecosystem of building a beautiful thing called Kubernetes—and it's really beautiful. It's really, really, very well thought off, and it's solving a lot of problem, and it's just making stuff extremely more easy—so if you look before that, and you needed to do stuff, like for instance, make sure that what's happening if your application is down, handling all your VMs and stuff, that is relatively way more a hard. Right now, basically, we’re abstracting to make it extremely simple. And with the managed solution in the Cloud like a AKS, GKE, and so on, it's become very, very simple.

So, I think that the main reason, to be honest, it's just, first of all, it's make you go faster, it's way more scalable, and the tooling is great. It’s just makes it so much simpler to do this and to do the other thing. So, that I think a lot of that reason, it's why you probably want to go to distributed, to use microservice [00:21:13 unintelligible] and monolithic application. And again, it's something you [00:21:16 unintelligible] problem. Like for instance, you can cut your organization into a small groups that basically can own a service, and it will be very seamlessly, kind of like, connected to each other, it's mean that each of them can write in a different language, they can write in—basically, own their own schedule for the versioning, and they don't really need to, per se, be basically consumed by the fact that they were working all on one big, giant source code.

I think all of this is very helpful. And then on top of it, with the fact that the tooling is very useful and make it very, very easy to work with, and your velocity of the team is going dramatically up. So, I say, I think that that's probably will be one of the reason for microservices and distributed application.

Emily: What do you think, among your customers, are some misconceptions that they have about distributed systems, about cloud-native in general?

Idit: I think, to be honest, I don't think that they have a lot of them. And the reason is because this ecosystem—don't forget that, I don't know, Kubernetes is relatively mature, Docker was a very long time ago. It's not that new anymore, so I think that there was so much marketing from all the big organization that I think that it's well understood of. I don't think that there is people who doesn't understand what microservices is. I don't think that there is any misconception there.

The only thing that I will say is that there is this notion of microservices is the best, and sometimes to be honest, microservices also a little bit more complex. I mean, it's way simpler to build it in one binary, and it's probably even quicker. I think that when you move into microservices, mainly, the thing that you need to think about is, if you just wanted to build a very simple application, probably monolithic will go faster and be simpler. If you’re using microservices, yes, you're getting a lot of benefit of the scale and so on, but it's also mean that you will need to handling as a distributed application, and therefore they care about different system for logging, and different system for how to communicate between these microservices, and a lot of other stuff that I think, yeah, it's a steep, curving, curving learn. So.

Emily: So, I wanted to ask if you wanted to add anything about, sort of, what you've observed about the market for cloud-native technologies, what end users are looking for, and then also what you've learned about running a company in this space?

Idit: Yeah, no. So, I mean, if you're looking at the market itself, I feel that as everything, as I said, is a journey. In the beginning, people were very excited about Docker, and then they move to the next level, we're all very excited about Kubernetes. And I think now that the buzzword we say is service mesh, and how to connect those microservices together. And I think there will be more and more stuff coming up.

So, I think that that's a—it's basically the evolution. We're always striving to be better. So, I think that's what I see in the market. As I said to you, we see a lot of type of customers in a lot of states on this journey, but in the [00:24:19 unintelligible], they all have the same journey? They're all interested in microservices. They're all very driven to get a lot of observability and security, and the ability to run between service—and microservice. This is why they will adopt service mesh, and we'll see what will come next after it.

Emily: Yeah. Just a couple more questions, which is—the next one, what is a tool, an engineering tool that you can't do your job without?

Idit: Huh. It's a good question. Yeah, I mean, right now, it's basically will be we're using a lot of [00:24:55 unintelligible]. Of course, behind the scenes, we’re using stuff like Docker and so on. I’m using a lot of command line. Like this is [laugh] the thing that I'm using a lot. Yeah, I'm using Visual Studio Code. That's something that is also extremely helpful for me.

Emily: Right. And then how can people connect with you or follow you?

Idit: So, they can connect me to Twitter, and my Tweet account is open, which is @idit_levine. You can also join our Slack. We have a huge community in the Slack. And so itself and all of us there. So, we are extremely active in answering any question from every reason. That could be another way to connect me. And the last one is to go to the website of the company and just send a message. That would be the last one.

Emily: Fabulous. Well, thank you so much for joining me, Idit.

Idit: Thank you so much for having me. It was a lot of fun.

Emily: Thanks for listening. I hope you’ve learned just a little bit more about The Business of Cloud Native. If you’d like to connect with me or learn more about my positioning services, look me up on LinkedIn: I’m Emily Omier—that’s O-M-I-E-R—or visit my website which is emilyomier.com. Thank you, and until next time.

Announcer: This has been a HumblePod production. Stay humble.

View Details

This conversation covers:

  • Mirage’s role as an API mocking library, the value that it offers for developers, and who can benefit from using it.
  • How Mirage empowers front end developers to create production-ready UIs as quickly as possible.
  • How Mirage evolved into an API mocking library
  • How Mirage differs from JSON Server
  • Sam’s relationship to Mirage, and how it fits in with his business. Sam also talks about open source business models, and whether Mirage could work as a SaaS offering.
  • One interesting use case for Mirage, which involves demoing software and driving sales.

Links

  • Mirage
  • Sam’s teaching site
  • Follow Sam on Twitter
  • Subscribe to Sam’s YouTube Channel

TranscriptEmily: Hi everyone. I’m Emily Omier, your host, and my day job is helping companies position themselves in the cloud-native ecosystem so that their product’s value is obvious to end-users. I started this podcast because organizations embark on the cloud naive journey for business reasons, but in general, the industry doesn’t talk about them. Instead, we talk a lot about technical reasons. I’m hoping that with this podcast, we focus more on the business goals and business motivations that lead organizations to adopt cloud-native and Kubernetes. I hope you’ll join me.

Emily: Welcome to the Business of Cloud Native. My name is Emily, I'm your host, and today I'm chatting with Sam Selikoff. Thank you so much for joining us, Sam.

Sam: Thanks for having me.

Emily: Yeah. So, today, we're going to do something a little bit different, and we're going to talk about positioning for open source projects. A lot of people talk about positioning for companies, which is also really important. And they don't always think about how positioning is important for open source. Open source maintainers often don't like to talk about marketing because you're not selling anything.

But you are asking people to give you their time which, at least for some people, is actually more valuable than their money. And that means you have to make a compelling case for why it's worth it to contribute to your project, and also why they should use it, why they should care about it? So, anyway, we're going to talk with Sam, about Mirage. But first, I should let you introduce yourself. Sam, thank you so much for joining me, and can you introduce yourself a little bit?

Sam: Sure. My name is Sam Selikoff. These days, I spend most of my time teaching people how to code in the form of videos on my YouTube channel, and my website, embermap.com. Most of it is front end web development focused. So, we focus on JavaScript. I have a business partner who also works with me. And then we also do custom app development, you know, some consulting throughout the year.

Emily: Cool. And then tell me a little bit about Mirage.

Sam: Yeah, so Mirage is the biggest open source project I've been a part of since falling into web development, I'd say about eight years ago, I got into open source pretty early on in programming, kind of what made me fall in love with web development and JavaScript. So, I was starting to help out and just get involved with existing projects and things that I was using. Eventually, I made my way to TED Talks, the conference company where I was a front end developer, and that's actually where I met my business partner, Ryan. And we were using Ember.js, which is a JavaScript framework, and we had lots of different apps at TED that were helping with various parts of publishing talks, and running conferences, and all that stuff.

And we were seeing some common setup code that we were using across all these apps to help us test them, and that's where Mirage came from. There was another project called Pretender, which helped you mock out servers so that you could test your front end against different server states. And we first wrapped that with something called Pretenderify, and then it grew in complexity. So, I was working on it on my learning Wednesdays, renamed it to Mirage, and then I've been working on it basically ever since. And then, the other big step, I guess, in the history is that originally was an Ember only project, and then last year, we worked on generalizing it so that it can be used by React developers, React Native developers, Vue developers, so now it's just a general-purpose JavaScript API mocking library.

Emily: So, we would say that the position is an API mocking library. And—does that sound right?

Sam: Yeah. If I had to say what it is, I would say it's a mocking library that helps front end developers mock out backend API's so that they can develop and test the user interfaces without having to rely on back end services.

Emily: Why does that matter?

Sam: It matters because back end services can be very complicated, there can be multiple back end services that need to run in order to support a UI, and if you're a front end developer, and you just want to make a change and see what the shopping cart looks like when it's empty. What does the shopping cart look like when there's one item? What does it look like when there's 100 items, and we have to have multiple pages? All three of those states correspond to different data in some back end service, usually in a database.

And so, for a front end developer, or anyone working on the user interface, really, it can be time-consuming and complex to put that actual server in that state that they need to help them develop the UI. That can involve anything from running, like, a Rails server on their computer to getting other API's that other teams manage into the state they need to develop the UI. So, Mirage lets them mock that out and basically have a fake server that they control and they can put into any state they need. So, it’s like a simplified version of back end services that the front end developer can control to help them develop and test the UI.

Emily: And when you first started Mirage, did you think of it as an API mocking library?

Sam: Not exactly. We used it mostly because of testing. So, in a test, it's usually a best practice to not have your test rely on an actual network. You want to be able to run your test suite of your user interface anywhere, let's say on an airplane or something like that. So, if your user interface relies on live back end services, that's usually where you would bring in a mocking library.

And then you would say, okay, when the user visits amazon.com/cart, normally, it would go try to fetch the items in your cart from a real server, but in the test, we're going to say, “Oh, when my app does that, let's just respond with zero items. And then in this next test, when my app does that, let's respond with three items.” So, that's the motivation originally, is in a testing environment, giving the UI developer control over that. And then what happened was that it was so useful, we started using it in development as well, just to help during normal times, just because it was faster than working with the real back end services.

Emily: Do you think there are any other projects that do something similar?

Sam: Yeah, for sure. I think the most popular one is called JSON Server, which is a popular open source library that lets a front end developer put some data in a file, and then you point JSON Server to it, and then it just gives you an instant mock server you can use to help develop your app.

Emily: So, what's the difference between Mirage and JSON Server?

Sam: The difference is that JSON Server is made—it's really optimized for giving you a development—kind of a fake server as fast as possible, but it comes with a certain format that it gives you the data in. So, what ends up happening is that it can help you get feedback and build your UI faster, but eventually, you're going to need to point your app at a real API server, whatever you planning on using in production. And so the way JSON Server works might not correspond—in fact, often doesn't correspond with your actual API. So, Mirage fills that gap because Mirage is designed to be able to faithfully reproduce any production API; there's ways to customize how the data comes back so that it matches so that as you're developing your actual user interface against Mirage, you can have confidence that it'll work once you switch over to production.

Emily: Is Mirage slower than other options?

Sam: Not performance-wise because they're all JavaScript code that runs in the browser, but JSON Server is really optimized for just getting started as fast as possible because it comes with all of those pre-baked conventions about how the data is going to be moving back and forth. So, with Mirage—it can be faster, it depends—but with Mirage, you need to learn a little bit more in order to understand how to faithfully reproduce your production API. But I think it's faster because in the long run, if you're writing code against a mock server that doesn't match the interface of your production API, then you're just going to be having to change that application code that you're writing.

Emily: How much do you talk to other people in the Mirage community, and talk about how they're actually using it?

Sam: I felt more in touch with the users when it was an Ember project only because Ember is a more niche-type community. Whereas now, there's folks using it in React and Vue, like I was saying, and Angular. And so, we get issues almost every day on the project. It's not like a mega-popular project, but it does have enough people using it that people will ask questions, or open an issue almost every single day. And so I try to stay in touch with the users through that, basically.

And then when I went to conferences—you know, before 2020—I would love to talk with people about it, or people would just bring it up. That's kind of how a lot of people know me on the internet. So, I would say that I do it in kind of a passive way. I haven't actively gone out to talk to them, partly because it's an open source project so it doesn't contribute to our revenue. We have some ideas for how that might happen one day, but as of right now, we can't justify doing proper product development on it in the sense of spending time doing customer interviews and stuff because it's a free project right now.

Emily: That is my next question, which is, how does it fit in with your business in the larger sense? You know, how you put food on the table? And what are your goals for the project?

Sam: A good question. I mean, it's really aligned with our overall mission of, of everything that we do because me and my business partner, Ryan, we really want to just help people get better at UI development, we want it to be easier, we want to help empower more people to do it because we think it's powerful tool in the world, and it's just too hard right now; there's so many things that are hard about it. So, that goes back to our consulting and our teaching; mainly our teaching. That's our main mission. And then Mirage is really—the purpose of Mirage is to enable front end developers to do more kind of with less, so they don't have to run a Docker container or get an SQL database up just to change some CSS for a given server state.

So, it fits into that mission. I've been doing open source long enough to know the pattern, and it happens over and over again, where people work on something, it gets popular enough that they start opening issues, and it becomes a maintenance burden for the maintainers, and then they try to stay up late closing issues, they get burned out, and then the project kind of rots. So, that's something that happens a lot in open source. And so, over the last five years or so of working on Mirage, I've been more involved, or sometimes step back if I need to spend more time on other things, but I'm really interested in making it sustainable, and I know some people in the open source community who have made their projects sustainable financially, other with a pro plan, or support plan, or different related services. So, if we could snap our fingers, that would be what would happen, and that's what we're working on now.

Emily: Which one of those open source business models do you think is most appropriate?

Sam: At this point, we basically are trying to not guess that answer because we think the way to find out which one it is, is to get a critical mass of users and listen to what they're saying. So, on our podcast, we interviewed Mike Perham from Sidekick, who runs Sidekiq and Sidekiq Pro in the Rails community for about 10 years, and he makes really good money. Sidekiq Pro came about because the enterprise customers of Sidekiq were asking for more robust job servers and all this kind of stuff, so there was a natural path for him to making a pro version that he could sell to enterprise clients. And so that's worked out really well for him, and the rest of the community gets to use the base version for free. And I love that because I do love this zero-cost to entry to open source.

And then my other friend, Adam Wathan who works on Tailwind, he's made money to help Tailwind be sustainable through education and courses, and then more recently, a project called Tailwind UI, which is pre-built UI components with Tailwind. And that, again, came about from people asking for that after he was working on Tailwind. So, I think the best way to do it is to get that critical mass of users to the point where you know what they're asking for, you hear over and over again, and then it makes sense to go forward with that.

Emily: There's also a third option, which is a cloud service, and I'm just curious—like a complete SaaS offering—have you ever considered that?

Sam: Absolutely. There's a really cool opportunity there for Mirage because your Mirage server usually lives alongside your front end code so that when every individual front end developer pulls the project down, starts working on it locally, they're running their own Mirage server. But some of the things that people have done organically with Mirage is, create a certain configuration of a server state—let's say for a demo—so let's say they create a shopping cart, or let's say they're working on a financial piece of software, and they need to show what it looks like with three clients and four contracts, and here are the products that we sold and how much money they are. People will make up a Mirage scenario with that specific set of data that their salespeople take and use on sales calls. And that way, again, the user interface looks fully realistic, it's a fully working UI that is talking to Mirage, but now you don't have to worry about the actual back end servers going down or anything like that.

And so, we've had this thought of a hosted Mirage, basically like a Mirage cloud, where even non-technical people could tweak the data there, and then again, they don't have to get involved with ops people or anything like that. And that could be really powerful. So, there's a ton of ideas there. I think our hesitation with our company, Embermap, it's worked to some extent, but it's not grown to the point where it can sustain us, and so that's partly because of the market issue, Ember being a little smaller. So, we're nervous to jump into any particular solution before we feel there's a proven market need for it. And so that's why, even though we have a lot of these fun ideas that I think could really work out, we first want to wait until we get that critical mass, the audience size that we'd feel comfortable could sustain a business.

Emily: How many people is that? What is the audience size?

Sam: That's a tough question. Let's say Mirage cloud was the goal. I think instead of waiting to a certain point, and then trying to build Mirage cloud, we would do a few things in the interim that would be little experiments and lower risk. So, I think the first thing would be, let's say, a Mirage course. And if we can sell a course on Mirage—or even we make a free Mirage course—that gets enough attention, that would tell us that there is enough of a need there, that the positioning resonates, that people are seeing it's a valuable thing, and that would tell us, okay, let's try the next thing.

So, we've been trying to do that. I make some YouTube videos over this last year, and I've been tweeting more about it and the work we've been doing, and so all of those are little experiments that I'm trying to pay attention to what resonates. Again, we have users who really love Mirage, and it's changed their workflow, but it's not at the point where I feel like it's a slam dunk and it makes sense to go on the next phase. So, I think there's been some frictions with Mirage, as we've brought it out to the wider JavaScript community. So, we have some things we want to tweak, and then ship a 1.0, and then I think maybe a 10 video course, would be a good next step, and just see the response to that, and basically take it from there.

Emily: Yeah. It's always interesting to think like, how will that you have reached the critical mass?

Sam: Yeah, I mean—

Emily: Hard.

Sam: It's really hard. And there's other strategies. I mean, a lot of businesses just have an idea and go try it out. There's this book I read, called Nail It and Scale It, and they talked about finding a problem. And we’ve have had some users of Mirage who use it at big, bigger companies, and it's really a big part of their workflow.

And we've thought about, “Hey, what if we were to build something for them that we could generalize?” They actually have a need for something like a Mirage cloud because every time they develop a feature, they have to sit down with their product people, and their front end developer will basically run through all these different Mirage scenarios to show them how the feature works in every case, and they would love to be able to just send a link to a hosted version where they could do that. So, we've thought about building that kind of thing.

But again, it's just a big risk. And then the market risk is there. So, what I've seen, I feel like, working in the past few years with a small circle of small business people and open source that I hang out with is this audience-first approach. So, it's like, if you're delivering a lot of value in the form of open source, and education, and talks, and you build an audience that believes what you have to say, and likes your opinion on things, likes your point of view, and again, resonates with the value of your work, then it becomes more and more obvious. You have a lot of people giving you feedback, you start seeing the same thing over and over.

And that's just what we do with developing Mirage itself. We know a lot of what we work on next is driven by the same issues that come up, the same problems that come up in the issues, over and over again. So, I think it is hard to know, but it's one of those lukewarm things. And basically, right now, it feels too early. You know?

Emily: And do you feel like, when you say to somebody, let's say, somebody who isn't involved with Mirage at the moment, maybe you're at a developer meetup or something, and you say, “I created Mirage. It's an API mocking library.” Do you have an aha moment? Are they like, “Oh, yeah. I know what that is. I know why you would use that.”

Sam: Mm-hm. Sometimes yes, and sometimes no. There is a form of position I feel like, sometimes, really resonates well with people, which is like, “Don't get blocked by your back end developers.” So, if you're a front end developer, and the API is not ready, you can still build your app, including all the dynamic parts of it which would normally require a real server to be running. So, you're the front end developer, and you're trying to wire up the interface, and what happens when you click save on a cart, and it saves it in the back end, and then you show a success message.

But your API is not ready, and you're working with another team that's working on the API, so you're kind of frustrated because you're stuck and you feel like you can't do anything, well, that's where Mirage can come in because you don't have to wait on them at all. You can just build it out yourself with Mirage. It's much simpler because it abstracts away all the complexity of a real server enough that you can actually build your UI against this kind of faithful reproduction of the back end. So, people really like that. And then people really like the testing use cases as well.

They get confused about the best way to test user interfaces in this new distributed world we're living in where you're maybe working on a React app, and the back end is a separate API that's also serving up iPhone clients and things like that. So, there's really multiple apps involved with running a website like amazon.com. So, the question is, how do you test the front end? And Mirage is a great answer for that as well. And that's where it originally came from, so.

Emily: Yeah. I can definitely see the positioning is sort of like a way for front end developers to decrease their dependence on their back end colleagues.

Sam: Yeah, exactly. It's hard because there's a lot that it helps you with. And the people who use Mirage the most and have used it the most, it's really transformed their workflow because it lets the front end developer just move so much faster. So, it's really just, like, the fastest way to build a front end. That's really what the whole point of—every change we make to the library, every decision that went into it, it's how to empower a front end developer to build a production-ready user interface as fast as possible. And that includes accounting for all these different server states that are usually a pain in the butt to get.

Emily: I'm just taking notes. I mean, that that actually sounded really powerful. Like, “Mirage lets front end developers work as fast as possible,” basically. It's fairly high level, but ultimately, that is a pretty compelling value statement.

Sam: Yeah. And I believe it to be true, too. And I mean, that's how—I mean, I've been working on apps recently, and I still use Mirage because it's just faster than using a real server because it's right there in your code. It's right alongside your front end code, so you don't have to switch over to another process running, or open up a browser and see it; it's all right there. If you want to switch from being an admin to being a user, it's like you just uncomment a line alongside the code that you're already writing.

So, I truly believe it is the best way to build a user interface. I think the positioning question is interesting because that is high-level. Sometimes developers in open source, they just want to know what it is, so there are some people who would see, “Just tell me what it is.” “It's an API mocking library.” “Okay, got it.” And that's what they want to know. “And what makes it different from other API mocking libraries?” Okay, we can talk about that.

But then there's other people I think, who could stand to benefit from it, and if they saw, “Mirage is an API mocking library,” they're going to be like, “I don't really need that.” But then they go back to their job, and they have to work on their app, and they find themselves spinning up a Docker container, going to auth0 to sign in just so they can run their app in an authenticated state, and it's like, Mirage would actually be perfect for this. You could just mock all that stuff out in Mirage once, all your front end developers could use it. And now anytime you want to just work on your app in an auth state, it's just right there in Mirage, you don’t have to deal with any of the complexities of the production services.

Emily: Yeah. I mean, I also like the idea of sort of the fastest way to build a front end app.

Sam: Yeah.

Emily: Or to build a UI. The fastest way to build a UI.

Sam: Yeah. That is, that's nice. I mean, it's interesting. It's a pretty interesting thing to say, and it's pretty compelling, and it's like, if you can back it up, that's a pretty strong claim.

Emily: Yeah. And I mean, one of the things that I also talk to clients about—so I work with all technical founders, usually, and we're often really focused on features, all the cool features, but—you don't have to ignore the features, but you sort of use them to prove that the thing that you're talking about, that the value you say you provide is not BS. But you still want to connect the dots. Like you were saying, sometimes even the most technical person is like, “Oh, a mocking library. But I don't need that.” They're not going to necessarily connect the dots that, like, “Oh, that's going to make my workflow way simpler.”

Sam: Right.

Emily: Does everybody know what a mocking library is? Like, everybody who needs one?

Sam: Mm, depends. If you're a front end developer. I mean, these days, people are becoming more and more specialized, so there's some front end developers who don't even deal with the data fetching side of an app at all. They're working on a part of the app that already has the data, and they're just working on maybe styling, layout, things like that, but they're never writing code or refactoring code that actually interacts with the server. And then even if you are doing that, you might not be writing tests.

So, a lot of people just do the lowest friction thing, which is, just point their local UI at whatever server they can. Maybe their company has a staging server or maybe they run one locally in development and they just build it like that. But again, that's the lowest friction, but it's slow because now you're running a Rails server. If you want to change the data, you have to know how to do it in Rails, and you have to run different commands. And then it comes to testing, it's really hard, too. So, I think most people would know, if you were to say, “Yeah, it mocks out the API.” Most people would know that.

But people might have different conceptions of what that means. So, there are some approaches to mocking—or stubbing, or faking, there's some technical difference between those terms, but for the purpose of this conversation, it's just faking out that functionality—there's some people who would hear that and say, “Oh, okay, I get it. A function that my code calls when it normally goes to the server, I'm going to replace it with a different function that returns this fake data.” But Mirage works differently because it operates at the boundary of the app. So, when you mock your API with Mirage, you don't have to change anything about your application because it intercepts the network request before it goes to the server.

So, that's another big benefit of Mirage, that Mirage has over other solutions for mocking out network functionality is that your application code stays exactly the same. And that's an important point because the whole goal with Mirage is to be the fastest way to build a production-ready user interface. And by production-ready, I mean an interface that can be plugged into your production API and you'll have confidence that it works. So, that boundary thing is another important point of Mirage. So, that would maybe be the one thing that people could mean different things when they say ‘mocking.’ But by and large, I would say most front end developers would understand if you said ‘API mocking,’ what they mean.

Emily: And do you think they generally are also able to figure out what that's used for, why they should care?

Sam: Yeah, that's the thing is that that's where I think the real opportunity is because I think, if you were to say ‘API mocking,’ people are going to immediately go to testing because that's the kind of environment where you would want to mock the API so you have control over it. But people who haven't used Mirage or something like Mirage, don't realize how powerful it can be, to mock it just during your normal development flow. So, again, once people use it and get it, they use Mirage for everything; it's just part of their workflow. I just start up my UI, I'm running Mirage in a specific scenario, and if I need to see the UI in a different state, I just changed my Mirage scenario, or put some new data in there, or empty out the database there. And so it totally affects your whole workflow because now you're just disconnected from the back end services.

So, I think a lot of people aren't doing that. They are just stuck—you know, they're just used to the hassle of getting these three services up and running just so that they can run their UI, and it's a pain in the butt. I mean, I was just talking to someone on another podcast, and he was saying it's the same thing. And it's really hard when they onboard new people because they have to get them set up with all these auth keys just so they can run the front end, even though they don't really care about the auth setup, they just want to start working on this page. So, again, the companies that have used Mirage and adopted it don't have any of those problems because the Mirage server is right there with the user interface, and they don't have to worry about it. So, I think that is a big gap that we could probably do a lot better job of closing with information on the homepage, or examples, or something like that.

Emily: So, Mirage also helps new hires get up to speed a lot faster, or get productive, I should say?

Sam: So, yeah. If you were hired by Facebook, and you're a front end developer, and your first task is to update the way the colors look on the home feed, to get that running on your computer so you can make changes and see how it looks in the code, at many companies—probably most companies—it's going to involve a lot of moving pieces because you have this React app on the front end, but it has to fetch data from somewhere so you can see how it works. Maybe they have a server that is just used for development that the person can point to. But again, if you have this shared hosted server, you only get what's on there; maybe you don't have an easy way to change what data it gives you. So, usually, front end developers need to see a dynamic user interface in all these different states.

But if you're using a shared server, you don't get all those different states. But it's easier to set up because it's already set up. So, the alternative is to say, “All right, Sam, you're new here. Let's get you set up so that you can run a copy of Facebook locally.”

So, that usually involves a ton of steps. I mean, I've seen that take, like, a week just to get all the pieces that are required to run the back end up so that you can actually power your front end. So, yeah, companies definitely, definitely have a problem with this. They have a problem with shared staging servers, I’ve talked to tons of people over the years who hate their staging servers, the back end teams, the ops teams hate them because everyone's always trying to change it for the salespeople to go take demo calls or for people to test. And so basically a shared server like that is just a nightmare because it's basically another production server to maintain that your internal team is trying to use and people want it to do different things. So, that's why Mirage is nice because it's just local to the code, and every front end developer can just tweak it in exactly the way they need for what they're doing at that time. So, I think that's a big opportunity as well.

Emily: One of the most interesting things, I think, about this conversation is that you've touched on this idea of salespeople doing demos, and I really like that because it's so different. It's such a different use case.

Sam: Yeah. We had a thought about that, actually because we were talking to a handful of companies—did interviews with them, actually, a couple years ago—and we're like, this could be a cool product. And it's like, just the easiest way to show off your software. And again, we've even done consulting for companies that have had basically another server, just for the purposes of a demo. And again, it's just a pain in the butt because it's another production server to manage, you have to reset the database after each call you do, and with Mirage is not like that; it just restarts with the app. It's just heavier, it's a real server to manage, where again if you just need to show off the UI, you can do it with Mirage. So, we thought about doing that. It's an interesting use case, for sure.

Emily: I think it's a really good illustration of how you can take the same technology, and reposition it, basically, to do something totally different, have a totally different set of value proposition. The people who would be actually benefiting from its use would be totally different. I mean, you have salespeople versus front end developers.

Sam: Yep. We've thought about that. What would that look like to market to those people? And maybe one way you could do it would be, you know, the salespeople get frustrated because they want to just change the numbers in this part of the app on the spreadsheet summary because they know it's going to help them communicate the value to the people that they're talking to, but now they have to go and ask some back end developer or someone on the ops team, “Hey, can you change this database column so that my demo works better?”

What if they had a Mirage-powered demo and they could tweak the Mirage data, maybe in some interface themselves, without having to involve anybody. And that's pretty compelling because Mirage, again, is designed to work with any API. So, no matter what your tech stack is on the back end, Mirage works for the front end team. And so now the back end team can do all their stuff, they can be switching from Rails to Elixir, they can be switching to Go or microservice architecture, and they have all this stuff going on, so they're the only ones who really know how to actually get this number from 100 to 200 in the system. Like, it's complicated. And so that’s, again, a benefit of just having Mirage because, on the sales calls, the people are just wanting to show what it looks like when the data is a certain way. So, yeah, that was definitely a use case that we didn't anticipate at all.

Emily: Yeah. And even you could have five salespeople trying to do calls at the same time and—

Sam: Exactly.

Emily: Wanting to have different data reflected, and I'm sure that's just like—

Sam: A nightmare. I mean, people have told us just, it's the worst. Basically, the shared staging server is a—yeah, it's a huge problem for a lot of teams.

Emily: It’s so interesting. Well, anyway, taking us back from the rabbit hole, although it—I mean, it really is interesting, and just fascinating how you could change this to be something that's targeted at that totally different market.

Sam: Right.

Emily: To wrap up. I mean, I was thinking about how you're saying that the fastest way to build a production-ready UI, that possibly is the most powerful thing I think you've said.

Sam: Yeah. We actually—I think that's how we had it when we first did the redesign of the site. And it was like, I don't know if that stuck in the sense of, what does production-ready mean? “Oh, I'm already writing production-ready code.” But then it's like, you have to—like you said, connect the dots and explain how you're not writing production-ready code unless the way you're mocking your API is matching your production API, so we kind of switched it to what it says now, which is, “Build complete front end features, even if your API doesn't exist.”

Which basically came from the mouth of someone who was using it, and when they had their aha moment that was what they were saying. I've been thinking if we get to the 1.0 launch, I think I would want to go back to something like that because I do think it's compelling. So, “It's the fastest way to develop a UI, test it, and then share a working demo of it,” because that's is really the motivation for the library. So, yeah, it might be interesting to think about doubling down on that as the unique value proposition.

Emily: Yeah. I'm curious what your immediate plans are.

Sam: So, we've been working on this REPL playground area of the site where people can learn Mirage and see examples, and we can share them and stuff. So, that will hopefully—you'll see a lot of that kind of thing in the JavaScript community. If you go to Svelte, which is another front end framework, you can create little sandboxes and play around with Svelte and learn it. So, it's a really good way to learn things and just see how it works quickly right in the browser without having to install anything. So, we're about wrapping that up.

And then I think it'll just be a matter of carving out the time to go through some of the issues and bugs that people have found, and getting it to a point where we feel good about slapping a 1.0 on it. At which point, we can take a look at all the feature requests and all the issues that people have run into and figure out okay, where do we want to focus our time? Again, considering, like, is there a story here where we can make it sustainable, and hopefully, dedicate more of our time to it? Because that's really, again, I think it's the biggest impact work I've done in my career so far, but it's just, sustainability in open source is tough.

So, it's a matter of not jumping too early on any particular idea for how to sustain it and monetize it: if it's a pro version, if it's a support plan, people have asked for all these things, but not maybe in the strongest numbers that would make me feel comfortable diving in on that. But I do believe that it can work because I believe it's a really good idea and I know a lot of companies have gotten a lot of value from it. So, that's what our short term plan is.

Emily: Well, fabulous. Thanks so much for talking about Mirage. This has been really interesting.

Sam: Yeah, thanks a lot for having me, and I appreciate your input. It was fun to revisit the positioning stuff. It's been a while, so it's always good to be thinking about that.

Emily: All right. I have one last question for you, which is what is an engineering tool you can't live without?

Sam: These days, it's got to be Tailwind. It's a styling framework. And it's just really excellent. So, I would not want to work on a site without it because it would just be painful. So, [laugh] I love Tailwind. Yeah, it's a really good library.

Emily: All right, cool. Well, thank you so much. And—oh, what—the very last thing it should be, how can listeners find you, follow you, keep up with you?

Sam: Yeah, best place is Twitter, at @samselikoff. And I'm also making YouTube videos more and more these days: youtube.com/samselikoff. So, those would be the places to go.

Emily: Right. Excellent.

Emily: Thanks for listening. I hope you’ve learned just a little bit more about The Business of Cloud Native. If you’d like to connect with me or learn more about my positioning services, look me up on LinkedIn: I’m Emily Omier—that’s O-M-I-E-R—or visit my website which is emilyomier.com. Thank you, and until next time.

Announcer: This has been a HumblePod production. Stay humble.

View Details

This conversation covers:

  • How Bloomberg is demystifying bond trading and pricing, and bringing transparency to financial markets through their various digital offerings.
  • Andrey’s role as CTO of compute architecture at Bloomberg, where he oversees research implementation of new compute related technologies to support kind of our business and engineering objectives.
  • Why factors like speed and reliability are integral to Bloomberg’s operations, and how they impact Bloomberg’s operations . Andrey also talks about how they impact his approach to technology, and why they use cloud-native technology.
  • How Andrey and his team use containers to scale and ensure reliability.
  • Why portability is important to Bloomberg’s applications.
  • Bloomberg’s journey to cloud-native.
  • Some of the open-source services that Andrey and his team are using at Bloomberg.
  • Unexpected challenges that Andrey has encountered at Bloomberg.
  • Primary business value that Bloomberg has experienced from their cloud-native transition.

Links

  • Bloomberg
  • Bloomberg GitHub
  • Follow Andrey on Twitter
  • Connect with Andrey on LinkedIn

TranscriptEmily: Hi everyone. I’m Emily Omier, your host, and my day job is helping companies position themselves in the cloud-native ecosystem so that their product’s value is obvious to end-users. I started this podcast because organizations embark on the cloud naive journey for business reasons, but in general, the industry doesn’t talk about them. Instead, we talk a lot about technical reasons. I’m hoping that with this podcast, we focus more on the business goals and business motivations that lead organizations to adopt cloud-native and Kubernetes. I hope you’ll join me.

Emily: Welcome to The Business of Cloud Native, I'm your host Emily Omier. And today I'm chatting with Andrey Rybka from Bloomberg, thank you so much for joining us, Andrey.

Andrey: Thank you for your invitation.

Emily: Course. So, first of all, can you tell us a little bit about yourself and about Bloomberg?

Andrey: Sure. So, I lead the secure computer architecture team, as the name suggests, in the CTO office. And our mission is to help with research implementation of new compute-related technologies to support our business and engineering objectives. But more specifically, we work on ways to faster provision, manage, and elastically scale compute infrastructure, as well as support rapid application development and delivery. And we also work on developing and articulating company’s compute strategic direction, which includes the compute storage middleware, and application technologists, and we also help us product owners for the specific offerings that we have in-house.

And as far as Bloomberg, so Bloomberg was founded in 1981 and it's got very large presence: about 325,000 Bloomberg subscribers in about 170 countries, about 20,000 employees, and more news reporters than The New York Times, Washington Post, and Chicago Tribune combined. And we have about 6000 plus software engineers, so pretty large team of very talented people, and we have quite a lot of data scientists and some specialized technologists. And some impressive, I guess, points is we run one of the largest private networks in the world, and we move about a hundred and twenty billion pieces of data from financial markets each day, with a peak of more than 10 million messages a second. We generate about 2 million news stories—and they're published every day—and then news content, we consuming from about 125,000 sources. And the platform allows and supports about 1 million messages, chats handled every day. So, it's very large and high-performance kind of deployment.

Emily: And can you tell me just a little bit more about the types of applications that Bloomberg is working on or that Bloomberg offers? Maybe not everybody is familiar with why people subscribe to Bloomberg, what the main value is. And I'm also curious how the different applications fit into that.

Andrey: The core product is Bloomberg Terminal, which is Software as a Service offering that is delivering diverse array of information of news and analytics to facilitate financial decision-making. And Bloomberg has been doing a lot of things that make financial markets quite a bit more transparent. The original platform helped to demystify a lot of bond trading and pricing. So, the Bloomberg Terminal is the core product, but there's a lot of products that are focused on the trading solutions, there is enterprise data distribution for market data and such, and there is a lot of verticals such as Bloomberg Media: that's bloomberg.com, TV, and radio, and news articles that are consumer-facing.

But also there is Bloomberg Law, which is offering for the attorneys, and there is other verticals like New Energy Finance, which helps with all the green energy and information that helps a lot to do with helping with climate change. And then there's Bloomberg Government, which is focused on, specifically, research around government-specific data feeds. And so in general, you've got finance, government, law, and new energy as the key solutions.

Emily: And how important is speed?

Andrey: It is extremely important because, well, first of all, obviously, for traders, although we're not in high-frequency game, we definitely want to deliver the news as fast as possible. We want to deliver actionable financial information as fast as possible, so definitely it is a major factor, but also not the only factor because there's other considerations like reliability and quality of service as well.

Emily: And then how does this translate to your approach to new technology in general? And then also, why did you think cloud-native might be a good technology to look into and to adopt?

Andrey: So, I guess if we define cloud-native, a little because I think there's different definitions; many people think of containers immediately. But I think that we need to think of outside of not just, I guess, containers, but I guess the container orchestration and scaling elastically, up and down. And those, I guess, primitives. So, when we originally started on our cloud-native journey, we had this problem of we were treating our machines as pets if you know the paradigm of pets versus cattle where pet is something that you care for, and there’s, like, literally the name for it, you take it to the vet if it gets sick. And when you use think of herd of cattle, there's many of them, and you can replace, and you have quite a lot of understanding of scalability with the herd versus pets.

So, we started moving towards that direction because we wanted to have more uniform infrastructure, more heterogeneous. And we started with VMs. So, we didn't necessarily jump to containers. And then we started thinking like, “Is VMs the right abstraction?” And for some workloads it is, but then in some cases, we started thinking, “Well, maybe we need something more lightweight.”

So, that's how we started looking at containers because you could provision them faster, and they could start off faster, and developers seem to be gravitating towards containers quite a bit because it's very easy to bootstrap your local dev environment with containers. And when you ship a container to the higher environment, it actually works. Used to be a problem where you developed on your local machine and you’d ship your code to production or higher environment, and it doesn't work because some dependency get missed. And that's where containers came about, to help with that problem.

Emily: And then how does that fit in with your core business needs?

Andrey: So, one of the big things is obviously, we need to ship products faster—and that's probably common to a lot of businesses—but we also want to ensure that we have highest availability possible, and that's where the containers help us to scale out our workloads and ensure that there's some resurrection happens with things like Kubernetes when something dies. And we also wanted to maximize our machine utilization. So, we have very large data centers and edge deployments—which I guess could be referenced as a private cloud—so we want to maximize utilization in our data centers. So, that's where virtualization and containers help quite a bit.

But also, we wanted to make sure our workloads are portable across all the environments, from private cloud to the public cloud, or the edge. And that's where containerized technologies could help quite a bit. Because not only you can have, let's say Kubernetes clusters on-prem on the edge, but also, now all the three major cloud providers support a managed Kubernetes offering. And in this case, you have basically highly portable deployments across all the clouds, private and public.

Emily: And why was that important?

Andrey: Basically, we wanted to have, more or less, very generic way to deploy something, an application, right? And if you think of containers, that's pretty much, like I say, Docker is pretty standard these days. And developers, we were challenged with different package formats. So, if you do any application Ruby and Rails, or Java and Python, there is a native packages that you can use to package your application, distribute it, but it's not as uniformly support it outside of Bloomberg or even across various deployment platforms.

But containers do get you that abstraction layer that helps you to basically build once and deploy many different targets in very uniform way. So, whether we do it on-premises, or to the edge, or to the public cloud, we can effectively use the same packaging mechanism. But not only for deployment, which is the one problem, but also for post-deployment. So, if we need to self-heal the workload. So, all those primitives are built there in the, I guess, Kubernetes fabric.

Emily: But why is being portable important? What does it give you? What advantage? I mean, I understand that's one of the advantages of containers. But why specifically for Bloomberg, why do you care? I mean, are you moving applications around between public cloud providers, and—

Andrey: So, we're definitely adopting public cloud quite a bit, but I guess what I was trying to hint is we have to support the private cloud deployments as our primary, I guess, delivery mechanism. But the edge deployments, when we actually deploy something closer to the customers, to your point about being faster, to deliver things faster to our customers, we have to deliver things to the edge, which is what I'm describing as something that is close to the customer. And then as far as the public clouds, we started moving a lot of workloads to public cloud, and that definitely required some rethinking of how do we want to adopt public cloud. But whether it's private or public, our main goal, I think, here is to make it easier for developers to package and deploy things and effectively, run faster, or deploy things faster, but also do it in more reliable way, right?

Because it used to be that we could deploy things to a particular target of machines, and we could do it relatively reliably, but there was no auto-healing, necessarily, in place there. So, resilience and reliability wasn't quite as good as what we get with Kubernetes. And what I mentioned before, machine utilization, or actually ability to elastically scale workloads, and—within vertical and horizontal—vertical, we generally knew how to do that. Although I think with containers and VMs, you can do it much better to higher degree, but also horizontally, obviously, this was pretty challenging to do before Kubernetes came about. You had to bootstrap, even your bunch of VMs and different availability zones, figure out how you're going to deploy to them, and it just wasn't quite there as far as automation and ease of use.

Emily: Let's change gears just a little bit to talk about a little bit of your journey to cloud-native. You mentioned that you started with VMs, and then you moved to containers. What time frame are we talking about? In addition to containers, what technologies do you use?

Andrey: So, yeah, I guess we started about eight years ago or so, with OpenStack as a primary virtualization platform, and if you look at github.com/bloomberg, you will see that we actually open-sourced our OpenStack distribution, so anyone can look and see if, potentially, they can benefit from that. And so OpenStack provided the VMs, basic storage, and some basic Infrastructure as a Service concepts. But then we also started getting into object storage, so there was a lot of investment made into S3 compatible storage, similar to how it AWS’s S3 object storage, or it's based on the [00:14:21 Sapth] open-source framework. So, that was our foundational blocks.

And then, very shortly thereafter, we started looking at Kubernetes to build a general-purpose Platform as a Service. Because effectively developers generally don't really want to manage virtual machines, they want to just write applications and deploy them to the—somewhere, right, but they don't really care that much about the where I use Red Hat, or Ubuntu or the don't really care to configure proxies or anything like that. So, we started rolling out general-purpose Platform as a Service based on Kubernetes. So, that was with initial alpha release of Kubernetes; we already started adopting it. And then thereafter, we also started looking into how we can leverage Kubernetes for data science platform. Well, now we have a world-class data science platform that allows data scientists to train and run inference on the various large clusters of compute with GPUs.

Then we quickly realized that on-prem, if we're building this on-prem, we need to have similar constructs to what you normally find in public cloud providers. So, as AWS add Identity Access Management, we started introducing that on-prem as well. But more importantly, we needed something that would be a discovery layer as a service, or if I'm looking for service, I need to go somewhere to look it up. And DNS was not necessarily the right construct, although it's certainly very important. So, we started looking into leveraging Consul as a primary discovery as a service. And that actually paid quite a lot of dividends and it's helped us quite a bit.

We also looked into Databases as a Service because everything that I described so far was really good for stateless workloads, to a greater degree with Kubernetes, I think you can get really good at running stateless workloads, but for something's stateful, I think you needed something, basically, that will not run necessarily with Kubernetes. So, that's where we started looking at offering more Database as a Service, which Bloomberg has been doing this quite a bit before that. We open-sourced our core relational database called Comdb2. But we also wanted to offer that for MySQL, for Postgres, and some other database flavors. So, I think we have a pretty decent offering right now, which offers variety of Databases as a Service, and I would argue that you can provision some of those databases faster than you can do it on AWS.

Emily: It sounds like—and correct me if I'm wrong, but it sounds like you've ended up building a lot of things in-house. I mean, you used Kubernetes, but you've also done a lot of in-house custom work.

Andrey: Right. So, custom, but with the principle of using open-source. Everything that I described actually has an open-source framework behind it. And this open first principle is something that now this is becoming more normal. Before we build something in-house, we looked at open-source frameworks, and we look at which open-source community we can leverage that has a lot of contributors, but also, can we contribute back, right?

So, we contribute back to a lot of open-source projects, like for example, Solr. So, we offer search as a service based on the Solr open-source framework. But also, we have Redis caching as a service, queuing as a service based on RabbitMQ. Kafka as a service, so distributed event streaming as a service. So, quite a few open-source frameworks. We're always thought of, “Can we start with something that's open-source, participate in the community, and contribute back?”

Emily: And tell me a little bit about what has gone really well? And also what has been possibly unexpectedly challenging, or even expectedly, but it's always even more interesting to hear what was surprising.

Andrey: I think, generally open-source first, as a strategy worked out pretty well. I think we have, I've listed only some of the services that we have in-house, but we certainly have quite a bit more. And the benefits, I think, I don't even know how to quantify it, but it certainly enabled us to go fast and deliver business value as soon as possible versus waiting for years before we build our alternative technology. And I think developer happiness also improved quite a bit because we started investing heavily into our developer experience and as a major effort. And this everything as a service makes it extremely easy for developers to deliver new products.

So, all of the investment we've made so far paid huge dividends. Challenges: I think that as with anything, starting with open-source projects, you certainly have bugs and things like that. So, in this case, we preferred to partner with a company that effectively has inside knowledge into the open-source project so we can have at least for a couple of years, somebody who can help us guide us and potentially, we—by actually invest in actual money into the project, we get it to the point where it's mature enough and actually meets a certain quality criteria. And some of the projects we invested heavily in which many people don't know, probably, but we—like Chromium Project. So, many people use Chrome, but Bloomberg has been sponsoring Chromium and WebKit open-source development quite a bit. JavaScript, V8 Engine, even the newer technologies like WebAssembly we’re heavily invested in sponsoring that.

But again, one thing that it's very clear, it's not just we're going to be the consumers of the open-source, but we're going to be contributors back with either our developers helping on the projects, or we need to invest to help this actual open-source project we're leveraging to be successful, and not just by saying, like, Bloomberg consumes it, but actually investing back. So, that's one of the things that was a big lesson learned. But currently, I think we have a really good enough system in place where we always adopt open-source projects in a very conscious and serious way with investment going back into the open-source community.

Emily: You mentioned being able to deliver business value sooner. What do you think are the primary two or three business values that you get from this cloud-native transition?

Andrey: So, ability to go faster. That's one thing that's very clear. Ability to elastically scale workloads, and ability to achieve uniformity of deployments across various environments: private, edge, public, so we are able to deliver now products to our customers as they transition to the public cloud, for example, much faster because we're have a lot of standardized and a lot of technologists that helped us with adoption. And including Kubernetes is one of them, but not only Kubernetes. We also use Terraform, extensively, some other multi-cloud frameworks.

And then also delivering things more reliably. That's I think, one of the things that is not always recognized, but I think reliability is a huge differentiator, and some of it has to do with how we deliver things to the customer with some resiliency and redundancy. So, we run very large private content delivery network as a service, and it's also based on open-source technologies. And the reliability is one of the main things that I would say we get from a lot of this technologies because if we do it on our own, yes, it would be generally Bloomberg working on this problem and solving it, but you get a, actually, a worldwide number of experts from different companies who’re contributing back to this technologies, and I see this as a, obviously, a huge benefit because it's not just Bloomberg working on solving some distributed system framework, but it's actually people worldwide working on this.

Emily: And would you say there's anything in moving to cloud-native that you would do differently?

Andrey: I think what I see as the big challenge, especially with Kubernetes, is adoption of stateful workloads because I still think it's not quite there yet. Generally, the way we're thinking right now is we leverage Kubernetes for our stateless workloads, but some stateful workloads require some cloud-native storage primitives to be there, and this is where I think it's still not quite mature. You can certainly leverage various vendors for that, but I really would like to see better support for stateful workloads in the open-source world. And definitely still looking for a project to partner with to deliver better stateful workloads on Kubernetes.

And I think, to a various degree, the public cloud providers, so hyperscalers are getting pretty good at this, but that is still private to them. So, whether it's Google, Amazon, or Azure, they deliver the statefulness to varying degree of reliability. But I would like this to be something that you can leverage on private clouds or anywhere else, and having it somewhere, well-supported through an open-source community would be, I think, hugely beneficial to quite a few people. So, Kubernetes, I think, is the right compute fabric for the future, but it still doesn't support some of the workload types that I would like to be there.

Emily: Are any other continuing challenges that you're working on, or problems that you haven't quite solved yet, either that you feel like you haven't solved, maybe internally that might be specific to you, or that you just feel like the community hasn't quite figured out yet?

Andrey: So, this whole idea of multi-cloud deployments, we leverage quite a few technologies from Terraform, to Vault, to Consul, to some other frameworks that help with some of it, but the day two alerting, monitoring, and troubleshooting with multi-cloud deployments is still not quite there. So, yes, you can solve it for one particular cloud provider, but as soon as you go to two, I think there's quite a few challenges that left unaddressed from just, like, single pane of glass—the view of all of your workloads, right. And that's definitely something that I would like to address: reliability, alerting across all the cloud providers, security across all the cloud providers. So, that's one of the challenges that I'm still working on—or, actually, quite a few people are working on at Bloomberg. As I said, we have 6000 plus talented engineers who are working on this.

Emily: Excellent. Anything else that you'd like to add?

Andrey: You know, I’m very excited about the future. I think this is almost like a compute renaissance. And it's really exciting to see all of these things that are happening, and I'm really excited about the future, I guess.

Emily: Fabulous. Just a couple more questions. First of all, is there a tool that you feel like you couldn't do your job without?

Andrey: Right. Yes, VI editor or [laugh] [00:27:42 unintelligible]? No. So, I think obviously, Docker has done quite a bit for the containerization. And I know, we're looking at alternatives to Docker at this point, but I do give Docker quite a bit of a credit because right now, local development environment, we bootstrap with Docker, we ship it as a deployment mechanism all over the place.

So, I would say Docker, Kubernetes, and the two primary ones, but I don't necessarily want to pick favorites. [laugh]. I really like a lot of HashiCorp tools, you know Terraform, Consul, Vault, fantastic tools; a really good community. I really like Jenkins. We run Jenkins, the service; really good. Kafka has been extremely reliable and scalability-wise, Kafka is just amazing. Cache in Redis is really one of my favorite cache tools.

There's probably a lot to mention. I've mentioned. [00:28:43 unintelligible] databases, Postgres is one of my favorite databases, or in so many varieties and different types of workload. But we also gain quite a lot from Hadoop and HBase. But one of my favorite NoSQL databases is Cassandra, an extremely reliable, and the replication across, I guess, low-quality bandwidth and of environment has been really awesome. So, I guess I'm not answering with just one but many tools, but I really like all of those tools.

Emily: Excellent. Okay, well, just the last question is, where could listeners connect with you or follow you?

Andrey: I am on Twitter as @andrey_rybka. I'm happy to get any direct messages. We are hiring. We're always hiring A lot of great opportunities. As I said, we’re open-source first company these days, and we definitely have a lot of exciting new projects. I haven't mentioned even probably 90 probably of other exciting projects that we have. We also have github.com/bloomberg, so you're welcome to browse and look at some of the cool open-source projects that we have as well.

Emily: Excellent. Cool. Well, thank you so much for joining me.

Andrey: Thank you very much.

Emily: Thanks for listening. I hope you’ve learned just a little bit more about The Business of Cloud Native. If you’d like to connect with me or learn more about my positioning services, look me up on LinkedIn: I’m Emily Omier—that’s O-M-I-E-R—or visit my website which is emilyomier.com. Thank you, and until next time.

Announcer: This has been a HumblePod production. Stay humble.

View Details

This conversation covers:

  • An average workday for Thomas as senior systems engineer at Systematic.
  • How Systematic uses cross-functional collaboration to solve problems and produce high quality software.
  • How security and data privacy relate to cloud-native technologies, and the challenges they present.
  • Systematic’s journey to cloud native, and why the company decided it was a good idea.
  • Why it’s important to consider the hidden costs and complexities of cloud-native before migrating.
  • What makes an application appropriate for the cloud, and some tips to help with making that decision.
  • The biggest surprises that Thomas has encountered when moving applications to cloud-native technology.
  • Thomas’s new book, Cloud Native Spring in Action, which is about designing and developing cloud-native applications using Spring Boot, Kubernetes, and other cloud-native technologies. Thomas also talks about who would benefit from his book.
  • Thomas’s background and experience using cloud-native technology.
  • The biggest misconceptions about cloud-native, according to Thomas.

Links

  • Systematic
  • Cloud Native Spring in Action book
  • Thomas Vitale personal website
  • Follow Thomas on Twitter
  • Connect with Thomas on LinkedIn

TranscriptEmily: Hi everyone. I’m Emily Omier, your host, and my day job is helping companies position themselves in the cloud-native ecosystem so that their product’s value is obvious to end-users. I started this podcast because organizations embark on the cloud naive journey for business reasons, but in general, the industry doesn’t talk about them. Instead, we talk a lot about technical reasons. I’m hoping that with this podcast, we focus more on the business goals and business motivations that lead organizations to adopt cloud-native and Kubernetes. I hope you’ll join me.

Emily: Welcome to The Business of Cloud Native. I'm your host, Emily Omier, and today I'm chatting with Thomas Vitale. Thomas, thanks so much for joining us.

Thomas: Hi, Emily. And thanks for having me on this podcast.

Emily: Of course. I just like to start by asking everyone to introduce themselves. So, Thomas, can you tell us a little bit about what you do and where you work, and how you actually spend your day?

Thomas: Yes, I work as a senior systems engineer at Systematic. That is a Danish company, where I design and develop software solutions in the healthcare sector. And I really like working with cloud-native technologies and, in particular, with Java frameworks, and with Kubernetes, and Docker. I'm particularly passionate about application security and data privacy. These are the two main things that I've been doing, also, in Systematic.

Emily: And can you tell me a little bit about what a normal workday looks like for you?

Thomas: That's a very interesting question. So, in my daily work, I work on features for our set of applications that are used in the healthcare sector. And I participate in requirements elicitation and goal clarification for all new features and new set of functionality that we'd like to introduce in our application. And I'm also involved in the deployment part, so I work on the full value stream, we could say. So, from the early design and development, and then deploying the result in production.

Emily: And to what extent, at Systematic, do you have a division between application developers and platform engineers, or however else you want to call them—DevOps teams?

Thomas: In my project, currently, we are going through what we can call as maybe a DevOps transformation, or cloud transformation because we started combining different responsibilities in the same team, so in a DevOps culture, where we have a full collaboration between people with different expertise, so not only developers but also operators, testers. And this is a very powerful collaboration because it means putting together different people in a team that can bring an idea to production in a very high-quality way because you have all the skills to actually address all the problems in advance, or to foresee, maybe, some difficulties, or how to better make a decision when there's different options because you have not only the point of view of a developer—so how is better the code—but also the effects that each option has in production because that is where the software will live. And that is the part that provides value to the customers.

And I think it's a very important part. When I first started being responsible, also, for the next part, after developing features, I feel like I really started growing in my professional career because suddenly, you approach problems in a totally different way. You have full awareness of how each piece of a system will behave in production. And I just think it's, it's awesome. It's really powerful. And quality-wise, it's a win-win situation.

Emily: And I wanted to ask also about security and data privacy that you mentioned being one of your interests. How do those two concepts relate to cloud-native technologies? And what are some of the challenges in being secure and managing data privacy specifically for cloud-native?

Thomas: I think in general, security has always been a critical concern that sometimes is not considered at the very beginning of the development process, and that's a mistake. So, the same thing should happen in a cloud-native project. Security should be a concern from day one. And the specific case of the Cloud: if we are moving from a more traditional system and more traditional infrastructure, we have a set of new challenges that have to be solved because especially if we are going with a public cloud, starting from an on-premise solution, we start having challenges about how to manage data.

So, from the data privacy point of view, we have—depending also on the country—different laws about how to manage data, and that is one of the critical concerns, I think, especially for organizations working in the healthcare domain, or finance—like banks. The data ownership and management can really differ depending on the domain. And in the Cloud, there's a risk if you're not managing your own infrastructure in specific cases. So, I think this is one of the aspects to consider when approaching a cloud-native migration: how your data should be managed, and if there is any law or particular regulation on how they should be managed.

Emily: Excellent. And can you actually tell me a little bit about Systematic’s journey to cloud-native and why the company decided that this was a good idea? What were some of the business goals in adopting things like Docker and Kubernetes?

Thomas: Going to the Cloud, I think is a successful decision when an organization has those problems that the cloud-native technologies attempt to solve. And some goals that are commonly addressed by cloud-native technologies are, for example, scalability. We gain a lot of possibilities to scale our applications, not only in terms of computational resources, and leveraging the elasticity of the Cloud, so that we can have computational units enabled only when needed. So, if there is, for example, an high workload on the application, and then scale down if it's not needed anymore, and that also results in cost optimization, but also scaling geographically.

So, with the Cloud, it’s more approachable to start a business that has a target in different countries and different continents because the Cloud lets you use different technologies and features to reach the users in the best way possible, ensuring performance and high availability. Something else related to that is resilience. Using something like Kubernetes and proper design practices in the applications, we can achieve resiliency in our infrastructure at a level that is not possible with traditional technologies. And then we have speed. Usually, a cloud-native transformation is accompanied by starting using practices like continuous delivery and DevOps, that really focus on automating and putting together different skills so that we can go faster in a more agile way and reduce the time to market. That's also a very important point for organizations.

Overall, I think there's a part of cost optimization, but at the same time, we should be careful because there are some hidden costs that sometimes are not considered fully. And that is about educating people to use new cloud-native technologies. We have some paradigm shift because some practices that were well-consolidated and used with traditional applications are now not used anymore, and it takes time to switch to a different point of view and acquire the skills required to operate cloud-native infrastructures and to design cloud-native applications. So, to make a decision about whether cloud-native migration is a good idea, I recommend to consider these hidden costs as well, not only the advantages but the hidden costs and the overall complexity, if you think about something like Kubernetes. For some applications, the Cloud is just not the right solution.

Emily: What would you say is a type of application that is not appropriate for the Cloud?

Thomas: As an application is not actively developed with new features, but it's in a pure maintenance phase and that is fully reaching the goals for its users. If it doesn't need to scale more than what it does today, if it doesn't need to be more resilient because maybe high availability is not that important or critical, then maybe going to the Cloud is not the right solution because you would add up complexity and make things actually harder to maintain. That is one scenario that I could think of.

Emily: So, basically, if it's not broken, don't try to fix it.

Thomas: Yes. So, cloud-native is the answer to a specific set of problems. So, if you don't have those problems, so maybe cloud-native is not for you because it's solving different problems.

Emily: So, would you say, the problems are if you need to really iterate quickly, develop new features, ship them out to customers?

Thomas: Yeah, if you don't need this agility because you're not actively developing features, or if it's small things that don't require so much speed or scalability power, then might not be a good idea. Or even if you don't have the resources, to acquire the skills required to design cloud-native applications and to manage cloud-native infrastructure.

Emily: What have been the biggest surprises as you've been using cloud-native technology and moving applications to cloud-native technology?

Thomas: So, I always think about how logging works. I think it's a fun anecdote that when you move to a cloud-native application, usually logging is managed through files in a traditional application where we set up rules to store those files. But usually in a cloud-native setting, like in a Kubernetes environment, we have the platform taking care of aggregating logs from different applications and systems. And these applications are providing these logs as events in the standard output. So, there's no files.

So, one of the first questions for developers when moving to the Cloud is like, “Okay, now where's the log files?” Because when something goes wrong, the first thing is, “Let me check the log file.” But there's no log file anymore, and this is just a fun aspect. But in general, I think there's a whole new way of thinking about applications, especially if we're talking about containerized applications. So, considering Java applications, for example, they're traditionally packaged in a way that needs to be deployed on an application server like Tomcat.

Now, we don't have that anymore, but we have a self-contained Java application packaged as JAR in a container. So, it's self-contained, it contains all the dependencies that are needed, and it can run on any environment where we have a container engine working, like Docker. So, developers take on more responsibilities than before because it's not only about the application itself, but it's also about the environment where the application needs to be deployed; that is actually part of the container now. So, we can see a flow of responsibilities that is different than before.

Emily: And do you think that that surprises a lot of organizations, that having developers need to take on more responsibilities is something that's not anticipated?

Thomas: Sometimes, maybe it's not anticipated. And usually, it's because when considering this migration, we don't consider those aspects of acquiring new skills, or bringing in, maybe, some consultants to help during the migration to help with all those practicalities, that from a high-level point of view, are not very visible.

Emily: Excellent. Tell me a little bit—I know you just wrote a book, so I was wondering if you could talk a little bit about your book and what inspired you to write it.

Thomas: The book is about designing and developing cloud-native applications using mainly Spring Boot, and Kubernetes, plus all the great cloud-native technologies that are available. The goal is to teach techniques that can be immediately applied to real projects, so to enterprise-grade applications, as much as possible. The cloud-native landscape is so complex, it’s so huge that it's impossible to cover everything, but I made a selection of all the aspects that I consider important and my goal is to try as much as possible to do things like I would do in a real application. So, what I would do daily in my job, so considering all those aspects that sometimes are not considered when teaching new technologies or new features, like for example, security. By the end of the book, the reader will have deployed a cloud-native system composed of different applications and services on a real Kubernetes cluster in a public cloud service.

And besides this, I also aim at navigating these cloud-native landscape because when we look at this famous landscape picture on the website from the Cloud Native Computing Foundation, it can be really overwhelming, both for new developers, but also for experienced developers that are experiencing different techniques—maybe, you come from a more traditional practice and want to switch to cloud-native, it can be really overwhelming. So, I hope that with this book, I can also help the reader navigating this landscape. The idea for the book, I got it back in January, February have been conducted by the Manning publisher, and about spring, we started considering different ideas. And I was really researching into the cloud-native landscape, in particular, how to use Spring and all related technologies to build strong native applications, so I decided to propose a book about that, that have these two main characteristics of teaching, as much as possible, real-world examples and techniques, and also help the reader navigating the landscape. So, it will not contain everything about cloud-native because that would require several books, but I think it's a good primer for the field.

Emily: Excellent. And just, actually, a couple, sort of, basic questions I forgot to ask at the beginning, like how long have you been working with cloud-native technology? When did Systematic start using it and moving applications to cloud-native?

Thomas: In Systematic, we have different projects and products. So, depending on that, some teams have started even several years ago, and in my project right now, it's quite recent. But I've been working with Spring for, I think, five years; also with Docker. And it's my favorite set of technologies because I think provides a lot of features to solve many different problems. I also, in my spare time when I have time, I like to contribute to the Spring projects on GitHub. It's a great community, I think.

Emily: Fabulous, and who do you think would benefit most from reading your book?

Thomas: I think it would benefit experienced developers, back end developers with experience in developing web applications in a more traditional way that would like to move to cloud-native, either because their organizations are doing the migration, or because maybe they would like to understand more how it works, or they would like to find a job in that field. But also for junior developers that have some experience with application development, so some basic experience with Spring and Spring Boot, but would like to take the next step towards the cloud-native world also understanding what is that about, and what, actually, cloud-native means.

Emily: What do you think is the biggest misconception about cloud-native?

Thomas: I think there's two misconceptions that usually, I find that is… thinking that cloud-native is about containers and that cloud-native is about microservices. So, for the first case, containers are used a lot for cloud-native applications, of course. They're used directly when working at the Kubernetes level, for example, but are used also when leveraging platforms like Heroku, or CloudFoundry. In that case, it's not the developer building the container, but it’s the platform itself, but still, they are used.

But cloud-native can also be applied to other different technologies, like functions, for example, in the serverless world. So, functions are not containers. So, in that way, I think it's a bit misleading to define cloud-native as containers. I think that containers is one of the technologies and implementations used for developing large native applications. But it's not the definition of cloud-native.

And microservices, also. I think it's wrong to imply that cloud-native means microservices because cloud-native applications are distributed systems. Of course, they can be microservices; it's a very used architectural style in the Cloud world, but I don't think that it’s a definition. Also, the CNCF previously had a definition for cloud-native technologies that was actually based on containers and microservices, and then they changed that. So, now they are listed as examples exactly because that is not the definition.

Emily: Fabulous. What about open-source, what do you think open-source and cloud-native’s relationship is?

Thomas: One example above all, like Kubernetes, that is open-source. It, I think, is the most popular project on GitHub with the highest number of contributions, if I'm not wrong. And I think that just explains everything about cloud-native because given that this is such a core project, and it's also the project that started all for the Cloud Native Computing Foundation, and it's a, now, very wide landscape of technologies. Open-source, I think it’s a very important part of it because it allows contributions from people around the world with different set of skills.

And we're not talking just about coding, but also other aspects, like, for example, technical documentation. I know that lately there have been several contributions, for example, to the technical documentation for Kubernetes to improve the documentations and help people approaching these technology and understand the more complex topics. And from a security point of view as well, from a testing point of view, I think that open-source technologies open many possibilities.

Emily: Great. Well, I just have a couple last questions for you. The first one I like to ask all of my guests, what is a engineering tool that you can't do your job without?

Thomas: That's a very difficult question because I use so many tools. But I will say if I had to choose one, I would probably say my terminal window.

Emily: Excellent. And then, how can listeners connect with you or follow you? And in fact, where should they go to buy your book?

Thomas: So, they can find me on my website, thomasvitale.com, where I also write blog post about Spring and security. I'm also active on Twitter. My handle is @vitalethomas and on LinkedIn. But on my website, they can find all my contacts there. I repeat, thomasvitale.com. And the book is available on manning.com. That is the name of the publisher, so they can find it there. The title of the book is Cloud Native Spring in Action: With Spring Boot and Kubernetes.

Emily: Excellent. Well, thank you so much, Thomas, for coming on the show and chatting.

Thomas: Thank you.

Emily: Thanks for listening. I hope you’ve learned just a little bit more about The Business of Cloud Native. If you’d like to connect with me or learn more about my positioning services, look me up on LinkedIn: I’m Emily Omier—that’s O-M-I-E-R—or visit my website which is emilyomier.com. Thank you, and until next time.

Announcer: This has been a HumblePod production. Stay humble.

View Details

This conversation covers:

  • The value that Forter provides, and the types of companies that they work with. Iftah also explains what makes Forter so unique.
  • The underlying technology that Forter is using, and how they quickly process hundreds of complex backend workflows. Iftah also talks about some of the tools that they are using, including AWS and Apache Storm.
  • How Forter approaches the cloud, and how it’s helping them concentrate on the business of detecting fraud. In addition, talks about the types of cloud services that Forter is using.
  • Forter’s ability to scale — including how they responded to increased customer demand during COVID-19.
  • Forter’s biggest technical challenge that they are currently working through.
  • Iftah’s thoughts on the security- speed tradeoff.

Links:

  • Forter
  • Forter on Twitter
  • Connect with Iftah on LinkedIn
  • Iftah’s email: iftah@forter.com

Transcript:
Emily: Hi everyone. I’m Emily Omier, your host, and my day job is helping companies position themselves in the cloud-native ecosystem so that their product’s value is obvious to end-users. I started this podcast because organizations embark on the cloud naive journey for business reasons, but in general, the industry doesn’t talk about them. Instead, we talk a lot about technical reasons. I’m hoping that with this podcast, we focus more on the business goals and business motivations that lead organizations to adopt cloud-native and Kubernetes. I hope you’ll join me.

Emily: Welcome to The Business of Cloud Native. I'm Emily Omier, your host, and today I'm chatting with Iftah Gideoni. Iftah is the CTO at Forter. Iftah, first of all, thank you so much for joining me.

Iftah: Very glad to be here.

Emily: So, I wanted to have you start by introducing yourself and what you do, and then also what Forter does.

Iftah: Hi, I'm Iftah. I’m a physicist of education, and in the last 20 years, a CTO of several companies, mostly [00:01:11 unintelligible] governmental companies, and companies that I founded. In the last six and a half years, I'm with Forter. And what Forter started to do from 2014 is to provide what was, at the time, very bold vision of fully automated, fully cloud-based decisions about whether to allow or decline e-commerce transactions.

Now, from that time we actually implemented and executed that, we decide very many more than 3 million transactions every day, today, all in real-time without a human in the loop. And we expanded into being a fully-fledged trust engine that gives decisions not only about transactions, but about many other points of interaction with the consumer, for example, in their login time, and in other points where trust decision is needed.

Emily: So, just because I think it might be interesting to listeners, give me some examples of, like, when somebody might interact with Forter or have some sort of action approved or declined by Forter.

Iftah: Right. The prime customers of Forter are the big e-commerce enterprises. Think about the [00:02:42 Sephoras], the Nordstroms, the Home Depots, and this kind of companies. And whenever you press the button of requesting to committing to the purchase and you see this small things rounding on the screen, then it is sent to Forter and Forter within, usually, half a second returns a decision.

Now, Forter does not act as an additional data point, or input, or score into some system of the merchant. It actually answer whether to approve or decline the transaction. In very many—and most of the revenue of Forter comes from a covered transaction that, if this transaction was fraud, it’s on Forter. Forter will guarantee it. And we were pioneering this model to putting our mouth where our money is.

Emily: Tell me just a little bit about why this is so difficult. What makes what Forter does unique?

Iftah: What Forter does is unique because it tells the human story, and takes it all the way to the decision itself. For example, it's very easy to approve the fourth transaction of a person that is sitting at home, browsing from home, making the purchase on the same desktop they made at previous times, and sending the shipment to the same home. That's very easy. But we want to be able to approve the traveler, the person that is sending a gift to a third party, or a person that is sending a gift to another state while not browsing from home and not from his common device.

We want to be able to approve those transactions that are checking out as guests from a new device and that's the first time this person ever appeared on our radar. And the ability to do that and to take the calculated risks and to look at the behavior, the cyber clues, and still be able to tell that this is indeed a new person and not someone that visited before and is trying now to hide. That's what makes what we do very difficult and complex.

Emily: So, tell me a bit about the technology story. What technology do you use to accomplish this, and how does it work? What does your stack look like?

Iftah: When I came to—from 2014, I looked at the system and what is actually needed in order to cater to such a complex story? And I thought to myself—and we'll talk about maybe a bit later about how all this is excellently suited for the Cloud, but what I found that throughput and big data is not the problem. First, it’s more or less solved, but it is the e-commerce business; it's not Facebook scale throughput. And on the other hand, it's not hardcore real-time, right? We're talking about tens of milliseconds, not the microseconds domain.

What is extreme about what we do is the complexity of the flow. We have hundreds of processes that are needed to be ran within that half a second in order to test, and check, and infer, and decide on many aspects of this transaction and of this person. So, first, we started from Amazon Web Services, and we started with, actually, Apache Storm. And why we decided that because we wanted to have something that enables first, a lot of parallelism—doing many things in parallel—with smart joins, that is with processes that takes information from other processes that executed in parallel, and can decide whether what they have so far from these processes is enough. Because we are very high availability, we didn't lose more than 10 seconds straight in the last four years. We are very high availability, but a lot of our sub-processes are not.

So, you need such a machine that will be able to infer about whether the information at hand is good enough and to move forward and still give, after half a second, the answer. We also wanted to have within this high availability system, we wanted to have the domain experts, the analysts, and the fraud researchers, we wanted to give them a very direct access to the code and each insight that they get, in close to real-time, maybe in 10 or 15 minutes from the time that they understood that there is a new wave of attacks or a new fraudster in action in a particular store or across stores. We wanted all these insights to be manifested in the system within 10 or 15 minutes without these people needing any engineering in order to do that. So, we created incubators within these Apache Storm processes that enable them to write, in Python, their wisdom into the system without being technologists or engineers. So, this was the basic.

Then we went on to see how we do the best similarity in the world. That is the understanding of whether we already saw a person, even if this person exhibits a new persona and is trying to hide. That is, they didn't give us the same phone number or email, it's not the same cookie, or the same IP, or the same credit card, and they don't use the same account, and we still want to know that this is the same person. This is a big part of what makes us efficient in exterminated fraud rings and enabling us to increase the cost of doing business for the sophisticated fraudsters. These are the prime building blocks and the last very important building block is the way we represent the world.

Usually and traditionally, world was represented by the transaction. The transaction was the building block. But we represent the world as people. We know more than 700 million people and of their interactions and their browsing, in many stores. These 700 million people include most of the people that interact online in the US. And the same is for the IPs, and the addresses, and the devices in the US. The US is where our coverage is best.

And all what we do revolves around the person because we believe that the person is what is actually persisting in the world with persistence reputation. That is a person is a legitimate person, they will stay legitimate. Usually, they won't flip on us. And if they are fraudsters, they will stay fraudsters. Not the same for IPs, for addresses, and for all other entities, you can think of. I hope this, it answered to a degree what you are asking about.

Emily: Yeah. And I'm going to go into some more questions, but it's it's really interesting that what you're combining is both this sophisticated technology as well as sort of an understanding of—almost like a law enforcement understanding of how fraud works. Or how—like, a anthropological investigation of how fraud rings work.

Iftah: Yes. And we found that there is a lot of—and I think our main asset is the ability to combine what analysts understand about the spoofing of the device, and how you detect that it's not really a mobile phone, it's an emulator on a desktop? And how can you tell that someone is trying to mess with an application that you protect? And what are the ways in which you can approve a transaction that looks very fishy to begin with, but it has some hints of legitimacy.

How we combine this with a very robust, high availability and very secure machine because it needs to be secure. We touch a lot of personal, identifiable information in our regular course of business, and we need the system to be ultra-secure while it is on the Cloud. And our booklet, actually, of 101, how to secure your startup [00:12:45 unintelligible] usage on the Cloud was actually trending number one on GitHub for months in 2017 when we issued it. [laughs].

Emily: That's excellent. Well, let me ask some more technology-specific questions. One is, just—you sort of alluded to this, but how is the Cloud important—and in fact, I believe you said critical to your business? Would Forter even be possible without the Cloud?

Iftah: Forter would be possible with an on-prem cloud, right, because when we say Cloud, it could be Amazon, or Azure, or GCP, but it also could be in a cloud that we built somewhere. This would be possible. We didn’t go there, and most of e-commerce companies would not go there, and we'll dive into this in a minute why it's not wise to go there. But Forter is heavily relying on knowledge of the people, regardless of which merchant they visited.

So, if we see a person in the Forter, it could be their first time Nordstrom sees them, but we already know them, and we can project the reputation of the person from previous interactions with other customers of ours. We don't share any customer data, of course, with any of our customers, but we can share parts of reputation, especially for people where this is the first time they visited a particular merchant. Now, this is a prime reason why it cannot be on-prem of the customer. And several customer—and I will not mention name, but huge conglomerates of carmakers, actually, asked us to be on their Cloud. And we refused and we let go of the business because that's not how we do, and the best value for them would be to share the data.

And so far, all the customers that we have so far actually agreed to share the usage of reputation with all the rest of the network of customers that we have. This is something that they cannot do in-house, and this is something, per your question, that cannot be done if we are not in the Cloud, but on their premises farm.

Emily: And so are you operating in all public clouds, or do you have your main technology running in one?

Iftah: We have our technology running in a few regions of AWS. And we are now deploying a few regions in Azure, too.

Emily: And so it doesn't matter if your customer, which public cloud. So, if you have a customer that uses GCP, doesn't matter, right?

Iftah: It doesn't matter. And most of them are [00:16:01 naturally] aware where we are. Bear in mind that we are serving companies—I mentioned the names—which are inherently not technology companies. And it doesn't matter where they sit; we are a full SaaS company for them. They send us the request, the transaction, and we give them a decision within this half a second, and that's the core of the business. Doesn't matter for them.

As [00:16:36 unintelligible] to say earlier, the concept of the public cloud and using other people's cloud infrastructure, be it GCP, or Azure, or AWS or others, is very suited for the e-commerce because of these two prime characteristics of the e-commerce: first, you don't need it to be very hard real-time, you're talking about tens of milliseconds, and giving answers in hundreds of milliseconds, ultimately; and second, unlike the Twitters, and the WhatsApps, and the Facebooks, and the Googles, the e-commerce is not big data in the sense that every transaction of e-commerce is, on average, a very high monetization. So, the ultra cost of using public cloud is definitely worth it for the e-commerce entity, comparing to creating your own farm. It is good for their flexibility and it's good for the focus and attention on their core business, where if you run your own farm, you are into a lot of domains of expertise which are far away from selling whatever you sell.

Emily: And tell me, also, a bit about, sort of, your own technology. Things like how you manage scalability. How important is it for Forter’s bottom line, the ability to have a scalable system?

Iftah: We are running from 2014—from day one, actually, from 2013, we have to be scaled out. We can’t scale up. We don't have anything that is done by a single computer. All the transactions are on what are called brains that are a scaled out on both redundancy and scalability.

All our data stores are scaled out. All the data stores that are storing the transactions, and the logging, and the entities that we talked about are scaled out and they are replicated, and the transactions that are dealing with our hundreds of thousands of browsing events that we receive and analyze every second, of course, they are scaled out. So, from day one, from 2014, everything that we do is scaled out. In the first two months, it was for redundancy in different availability zones of different regions, but from then on, it's all scaled out. And I will be very happy to dive into the particulars of the technologies, but what is important in the context of this podcast, I believe, is that doing it on the Cloud using the cloud infrastructure is actually enabling us to concentrate on the business of detecting fraud and business of these massive topologies of hundreds of processes that are both in Java and Kotlin and Python, and have very complex acyclic graphs connecting them. And we just do it in parallel on very many servers that we can scale up and out as we wish. And this is something that helped us focus our core business: understanding fraud.

Emily: Going back, actually, to this idea of scalability, I know over the past six, eight months because of COVID, e-commerce has gone through the roof. And I'm assuming, in fact, I read that Forter’s business has also been going through the roof. How have you managed scaling?

Iftah: Yes. Forter business went through the roof with their several verticals: with food deliveries, of course; and we the big department stores, which COVID accelerated their digital transformation; it did dive with the travel business, of course, right? Few things happen to our customers, and for Forter, scaling was natural. If we have 20 minutes warning scaling is a natural to us, and here we had about 10 days of warning. Easy, right?

For our customers, it was a bit different. First, a lot of them came to us, actually had to eliminate all their manual processes. And Forter was there for them. Now, what Forter did for them beyond eliminating any manual fraud-related tasks and loads was to reduce substantially the hike in the customer success load. Because Forter is able to be more accurate and to decline less legitimate customers, you don't have that many calls to the customer success centers.

And these are two bottlenecks for our big merchants: the customer success, and fraud and fulfillment. And the fulfillment, that is being able to capture the money and to send the goods is also streamlined by the fact that it's all done in real time. These are the direct effect, but there are additional phenomenon that happened. One of them is that suddenly, in COVID, a lot of customers that didn't usually do things online started to buy online. And we saw that the amount—or the percentage of new buyers, in many of our customers, suddenly jumped.

And when you have new buyers, you need a very sophisticated system to be able to allow them in, to approve their transaction, and to allow them to build their reputation; so this happened. The spikes in throughput happened; every day in the last four months is like a Good Friday and Cyber Monday combined for us. And that's good. We didn't lose any availability, and with the current technology, it wasn't that problem from the scalability aspects. And indeed, have we been on private, or our on-prem, this would be much harder.

Emily: And now tell me a little bit more about the technology required. We talked a little bit about it not being exactly a throughput problem, but you do have hundreds of processes that you have to run in, you know, several seconds, what technology do you need to leverage in order to make that happen?

Iftah: Everything that we run is on flash disks; we don't have rotating disks anymore. We do run low CPU and low memory on all—low memory usage and low percentage of CPU on every [00:24:30 unintelligible] that we run, to accommodate and reduce a spike pickups. We do use the Apache Storm for our base; it is the base of our topology of these processes that we talked about, and hundreds of them in each topology. And we have several topologies for both the transaction time and what we call the visit time, the browsing time. And we run—we are a big customer of Elasticsearch, we run Elastic from the very early days, and we use them for sophisticated queries in their own annoying language. [laughs].

And we have one of the largest clusters. We have about 15 clusters of Elasticsearch that serve our entities that we talked about, our mapping of people to these entities, our logging, and our real-time matching between the current transaction and all the hundreds of millions of people we already know that acted online previously. These are the core technologies in our stack, and on top of that, we use several other technologies: Spark and our wrappers over Spark for the MapReduce work of our machine learning processes, and we use Kafka for persistent distribution of our data among regions, and among availability zones.

Emily: What would you say is your biggest technical challenge? And by this, I mean, like, something that you're perhaps still working on, you don't feel like you've totally figured it out yet.

Iftah: I think we are very advanced in our matching, the similarity problem. That is something that we think is a pillar of our superiority in this field, but it's a never-ending story. The ability to detect relevant anomalies in the behavior of the crowd is something that we work very hard on, and we expect a lot from these technologies because they have the potential to help us mitigate threats which are new to us; zero-hour threats of modus operandi, of MOs that we did not encounter earlier. These are the main issues.

One issue that is mundane and prosaic is the cost of transaction. We do a lot of processing and we start, in our scale, to feel the heat of the cost of serving all these transactions. Nothing that will take us out of the Cloud, but it's something that we need to work hard on. Last, but definitely not least, is security. We think we turned our emphasis on security to our unfair advantage in this field, but still, hardening your systems and thinking about the possible attack vectors on your systems and on your merchant’s system is something that I lose sleep at night over, and it is something that we can never say that we are done with.

Emily: What do you think about the security-speed trade-off? Do you think it's real? Or do you think you can move just as fast and be secure?

Iftah: We can move with negligible sacrifices for the security. Again, if you are talking about real-time systems where the microseconds count, then it's a different story. But for us, having everything encrypted both at rest and at motion is something that does not need to come at the expense of security. What is very interesting in this trade-offs of security and speed is the trade off, not of the real-time speed and the processing speed, but of the engineering development speed. And here, the magic is in the automation, every security aspect, and with your ability to mask all these security aspects from your engineers, and giving them the right APIs so they can develop the application itself. Which is developed to our domain in the same speed, while still being totally secured without them needing to take care of the plumbing. And that's something that we invested a lot in, and it's a never-ending game. I think we're good at it, but never good enough.

Emily: And do you rely primarily on the Cloud service providers? So, on AWS’s native services, or do you tend to find additional out-of-the-box services, or build your own? Do you have, sort of, a philosophy on that?

Iftah: You know, philosophy is one thing, and then what you're doing practice sometimes need to be traded off with reality. But we are currently running on both AWS and starting to run on Azure, so we are making our processes agnostic to the particular cloud that we run on. It is interesting to do when you come to security configuration because you need to create abstraction layers over the particular security mechanisms in AWS and Azure, which are quite different. And that's where we are now. So, we are moving to be totally agnostic. So far, we did use occasionally, not—we weren't a heavy users of AWS services, but we did use a analytic databases; we did use Kinesis, but we moved now to Kafka, and so on. And we did use very cloud-specific queues, but we're moving out of this now.

Emily: Why do you think it's important to be cloud-agnostic?

Iftah: Because we run on two different clouds. We run on two clouds because of the very high availability requirement that we have. First, we need to be totally available to our merchants. Second, we need not only to be totally available to merchants, we also need to be very, very accurate, always.

So, it's not that I can degrade gracefully and say, “Okay, I always answer approve in certain occasions,” because the fraudsters will very quickly understand that. So, we need to be with full brain capacity, always on. And if we are not, we started within tens of minutes or a hour or two, to be very susceptible to great losses. So, that's the reason we need to be with multiple regions, and we need to be with both clouds. It does take a heavy penalty, and we do think about how to reduce the penalty of working with two clouds, but that's what we currently do.

Emily: Can you tell me how much your technology stack has changed since 2014?

Iftah: We did change a lot in the representation of the world, and this was big. We did move into a Elastic from more traditional NoSQL and SQL [00:33:20 unintelligible] BMS. And we move now, again, to new high throughput databases for our browsing events, the ones that do get hundreds of thousands of events per second. And we do move slowly [00:33:40 unintelligible] many more items or small stack items like queues, and data distribution channels that are no longer serving us well as we scale out, and as we move to being cloud-agnostic. For example, we move now our analytics database from AWS’s Redshift to a cloud-agnostic database.

Emily: Excellent. I'm going to wrap up pretty soon; this has been really interesting. But a couple, sort of, last questions I wanted to ask. One is, can you describe what a day looks like for you? What does the day for the CTO of a Forter, of a cloud-based SaaS fraud prevention company—what do you actually do?

Iftah: First, I am looking at what may endanger our business in the next year and in the next three years. The reason why we are [00:34:42 unintelligible], we call this process internally, the ‘what can kill us?’ process. Is mainly because we are in a good shape, and when you're in a good shape you need to look at the threats and how to protect the business from them, and what new business you need to do.

Then I'm looking at the health of our precision teams. And our precision teams are both the data science team, the cyber R&D teams, the fraud researchers teams, and the engineering teams that are supporting them. All these are—we need to see that we maintain our superiority. We so far never lost a QC or a bakeoff on any performance issues, and it's a tall order to keep it that way. So, this is the second task that keeps me up.

And the last is to see that, indeed we have all what we need in order to enable the spear of development and for the new products. Companies in their seventh year, as we are, are in an inflection point between the startup and the enterprise, and that's where you need to make sure that we stay agile. We stay agile, it depends on the agility of the organization. How can you scale? Or do you rely on several heroes? And the agility of the development itself that relies, to a great degree, on the tech debt kept low enough.

Emily: Fabulous. And what is a tool or platform that you think is sort of essential to functioning?

Iftah: I think that we built a very robust, extensive monitoring and alerting infrastructure, and this monitoring and alerting infrastructure enables us to understand quickly whether something has happened in the world. And I must say that most of the time that something is happening in the world, it's not something that we need to do something about manually, but sometimes it's something that the merchants need to do. We discovered, for example, that one of our online travel agencies customers started to issue flight tickets for one percent of their price; for three and four bucks instead of four hundred bucks. And we detected it not by looking at the prices, but by seeing spikes of purchases from this OTA in Malaysia and Vietnam, and we were able to tell this to our merchant, to the customer, and the whole thing was rectified about 14 minutes from the time it started. So, our alerting and monitoring systems, which is both on the application level, on the business level on them, and on the other end, on the machine levels, this is very, very important, and pays for itself handsomely.

Emily: I think I've read accounts in the newspaper of travel agents, or airlines having that type of mistake.

Iftah: Yes.

Emily: It tends to get some publicity. Last question is, how can listeners connect with you or follow you?

Iftah: Well, iftah@forter.com. Look us up in forter.com, and we will be very happy to talk to you.

Emily: Excellent. Thank you so much, Iftah, this was really fascinating.

Iftah: Thank you very much for having me.

Emily: Thanks for listening. I hope you’ve learned just a little bit more about The Business of Cloud Native. If you’d like to connect with me or learn more about my positioning services, look me up on LinkedIn: I’m Emily Omier—that’s O-M-I-E-R—or visit my website which is emilyomier.com. Thank you, and until next time.

Announcer: This has been a HumblePod production. Stay humble.

View Details

This conversation covers:

  • Laying the groundwork for a successful open-source program office (OSPO).
  • Why legal and engineering are usually the two main stakeholders in open-source projects.
  • Why engineering teams tend to struggle at articulating their perspective on open-source. Tobie offers some improvement tips.
  • How Tobie defines open-source strategy. Tobie also explains the risk of not having an open-source strategy, as well as his process for helping organizations determine the best strategy for their needs.
  • Common challenges that businesses face when deploying open-source software.
  • The secondary — or non-code — benefits of open-source, and why many organizations tend to overlook them.
  • Tips for engineers in non-technology organizations like pharmaceuticals or finance to approach business leadership about open-source.

Links

  • UnlockOpen: https://unlockopen.com/
  • Twitter: https://twitter.com/tobie

TranscriptEmily: Hi everyone. I’m Emily Omier, your host, and my day job is helping companies position themselves in the cloud-native ecosystem so that their product’s value is obvious to end-users. I started this podcast because organizations embark on the cloud naive journey for business reasons, but in general, the industry doesn’t talk about them. Instead, we talk a lot about technical reasons. I’m hoping that with this podcast, we focus more on the business goals and business motivations that lead organizations to adopt cloud-native and Kubernetes. I hope you’ll join me.

Emily: Welcome to The Business of Cloud Native. Today, I am talking with Tobie Langel from UnlockOpen, and I wanted to start, Tobie, by just asking, you know, what do you do? Can you give us sort of an introduction to what you do, and how you tend to spend your days?

Tobie: Sure. So, I've been back into consulting for a number of years at this point. And I essentially focus on helping organizations align their open-source strategy with business goals. So, it can be both at the project level—so sometimes helping specific projects out—or larger strategy at the corporate level.

Emily: So, I actually recently had Nithya Ruff, who's the head of the OSPO at Comcast on the podcast. For listeners who don't know, that's an open-source program office. So, are you sort of an outsourced OSPO for companies that aren't Comcast’s size?

Tobie: So, that's a really good question. My answer would be no, but it tends to happen that I help companies build that capacity internally. So, I would generally tend to come up before an OSPO is needed, and help them figure out what exactly they need to build. For OSPO, my pet peeve is companies building OSPOs like they need to tick a checkbox on the list of the things that they have to do to be up-to-date with good engineering practices, if you will.

In general, if you want to be successful, with an OSPO, it has to meet the particular needs of your company, and that's usually kind of hard to figure out if you just leave it to whoever in the organization is more interested in driving that effort. And so essentially, I sort of help in the early stages of that by bringing all of the stakeholders at the table, and essentially listening to them and making sure that what they want out of an OSPO is aligned between the different stakeholders and matches the overall strategy of the company.

Emily: And who are the stakeholders that you're generally talking to?

Tobie: So, essentially, open-sources is strange, for one reason, in terms of how it was adopted in companies from a historical perspective. Adopters have always been essentially engineers who just wanted better tools, or the package or the software that best fitted their current intention, and there's a very, very grassroots process by which companies start using open-source. And what happened at some point is companies sorted to see all of the software, and got concerned, and started trying to assess the risk. And so companies just tended to bring in the legal arm and lawyers at this point. And so to fulfill compliance questions, you bring in lawyers, and then the responsibility of grown-up open-source kind of falls on to lawyers, which tends to be problematic from the perspective of good engineering practice and velocity that you want from your engineering and product side in a company.

And so clearly, the two stakeholders or the two main stakeholders tend to be legal and engineering, and there tends to be a tension between these two sides. And in lots of companies this tension, instead of being resolved to some degree, tends to be won by the legal side that understands business concerns better and is better able to praise or explain what they do in terms of business impact and business risks than the engineering side. And so this equilibrium tends to create OSPOs which are legal heavy, process heavy, and don't really give engineers the kind of freedom that they would need to be effective in their daily engineering practice. And the reason behind that being essentially over exaggerated risk perception of open-source because, to be frank, open-source is not well taught in legal school and clearly not part of the curricular that most lawyers are familiar with when they move into helping tech companies out. So, essentially, I sort of tried to bridge these two worlds.

Emily: I can imagine that being an open-source lawyer, that's a niche, that's a very specific niche.

Tobie: Yeah, actually there's a running joke in that community, which is, “As soon as you get your law degree and you’re an open-source lawyer, you’re one of the 25 best open-source lawyers in the world.”

Emily: [laughs]. That's awesome. Why do you think engineering teams are so bad at clearly articulating their perspective on open-source, and what can they do to improve?

Tobie: So, there are clearly multiple reasons why engineers aren't the best at articulating how open-source matters. So, I think one of the key ones, it's just, it's something that's part of their daily practice, and they don't really understand and never have been taught the actual intellectual property—IP—impact, that open-source has on their company, and they don't really understand how others in the company might perceive this IP impact. So, I think, one part of it is, essentially, this is just how engineers work. Like, you want to use a piece of software, you put it in it, right? If you want to fix something, well, you do a pull request. This is sort of, like, a common practice. And it's always hard to articulate things that are essentially part of your, like—you know, like a native language, like part of your culture. It's really hard to describe, why you would do this, and why it matters. So, I think that's one reason.

The other reason, I think, is that there is a lot of overlap between the way legal works, and the way business works in general. Few examples of that are, engineers tend to think really like in binary way, like, you know, something is true or false, something is on or off, whereas business and law a much more spectrum thinking and into the gray area of things. Similarly, law will share with executive manager’s schedule, versus a maker’s schedule. So, there's lots of cultural artifacts of law culture in corporations that are much closer to business culture, and so, just a better understanding. So, I don't think engineering is really bad, per se. I think it's just bad when you compare it to legal, essentially.

Emily: I mean, and clearly, like, lawyers, their whole training is about making arguments for things that they believe to be true. So.

Tobie: Fair enough, but honestly, when you hear engineers talking to one another, that could be said, of engineers, too.

Emily: That's fair.

Tobie: Your second question was, how can engineers improve that?

Emily: Yes.

Tobie: And I think that's actually something that they can do and that has way more benefits than just making it easier for them to contribute to open-source, or to have a strong open-source culture at their company. And I think that's essentially focusing on the customer-facing business value, if you will, of what they are building. And if you can start articulating all of what you do in terms of how it affects the business, how it affects end-users or end-customers of your products, it gets way easier to have weight in conversations with other people within an organization that reason about this that way.

Emily: And I would imagine this applies not just to making a case for open-source, but everything in engineering. Making a case for using containers, making a case for changing something in your architecture, investing in engineering, hiring a new person—

Tobie: Absolutely.

Emily: —you have to learn to make the case in terms of the business impact.

Tobie: Yeah. It's interesting because we always look as growing up or leveling up as an engineer in terms of actual ability in your craft. But what really makes a difference is how you can leverage your craft to pursue broader goals, organizational goals. And yeah, you're absolutely right that skill set is useful, just, like, across the board. So, are soft skills, by the way, which is another thing that engineering tends to forget about, unfortunately.

Emily: So, going back to what you do in crafting open-source strategies, what is an open-source strategy, and what's the risk of not having one?

Tobie: So, by strategy, I sort of think about the plan that you have to meet certain goals that you care about meeting. And so an open-source strategy can be widely different depending on what those goals are, and what those organizational goals are. Some companies will have—their main business will be extremely tied to open-source software—you know, think like a company like MongoDB, or Redis, or Mozilla, for example—but for most companies, their business is kind of far away from actually producing open-source software. And so, an open-source strategy for those will be one that is more aligned with, like, how exactly can open-source help our organizations serve our clients better? The same way you would use DevOps to some degree. Or even, like, you know, Cloud, for lack of a better example. So, really, about how can you leverage these tools to help meet organizational goals?

Emily: And then what happens if you don't have a strategy?

Tobie: Oh, well, what—that's what happens when you're missing a strategy for anything else: you essentially end up at best copying what others are doing—so, you know, you're sort of late to the game—and that worse, just running around aimlessly. If you don't know why you're doing something, you don't know what to measure. And this is true of everything. I mean, this has nothing to do with open-source. You don't know what to measure, you don't know where to invest, you don't know if what you’re investing is actually giving you a useful return on investment. You know nothing, and so you're probably better off just not doing anything.

Emily: When you meet with these different stakeholders in a company, how do you help them figure out what the best strategy for that particular company is going to be, in relation to open-source?

Tobie: So, if we're looking at companies who are not essentially trying to monetize an open-source project, the way I usually start looking at that is looking at what are the current points of frictions? What are the challenges and the problems that a company is facing to run its software, its engineering operations with the kind of performance level that it would want to do? And this can be broadly different things. It could be an organization finds itselfs to be fairly siloed, and finds it really hard to collaborate with teams in different parts of the organization. It could be having a really hard time filling in their hiring pipeline, or having retention issues. There are just plenty of different problems that show up.

Then the second thing that we tend to look at is if they had a magic wand, if you will, what would their future look like? What would they want to achieve? And once we have this current situation and future desired state, we look to see at what part of open-source can actually help this transformation. And for that, what I do is I—there's a talk that I've given a number of time, called, “Making the Business Case for Open-Source,” which essentially focuses on all of these different aspects of open-source that are beneficial to companies, which I called byproducts, or second-order benefits of open-source, which is not the output of the code itself, but all of the benefits that having a strong open-source culture brings to a company. And we'll look at those, and we see if there's a good fit.

Emily: And how aware do you find business leaders to be about the secondary benefits of open-source? The sort of non-code benefits of open-source?

Tobie: Mostly not, honestly. I guess it's actually surprising how few companies get that, outside of the tech giants, by the way. All of the large tech companies understand that really well. Everyone from Google to Microsoft to Facebook to Mozilla, everyone is doubling down on these aspects and knows that open-source is where you tend to find a lot of really good engineers and that open-source really benefits engineers and helps them level up, and helps them build things that are actually, then—end up being really useful internally, like soft skills.

I mean, I know that open-source has a really bad rap, and there are reasons for this, and there are lots of things that, as a community in open-source, we have to improve. I don't want to be dismissive of that at all, but if you’re actually able to collaborate and get alignment in a large open-source project will you have—you can't go through like your manager to get your manager to speak to the manager of the person that's not complying with whatever it is that you want because it turns out, they're in a completely different company. When you're able to be effective in an environment that is as hostile as that one, once you bring that skill set back internally, you're highly effective. So, these benefits exist, and large tech companies understand these benefits really well. Outside of tech, though, that's not the case.

And when you look at the data, it's that's really telling because we have today really good datasets, per industry, of how much different industries use open-source, and frankly, at this point, pretty much there's open-source everywhere, in every industry, and in every project in every industry. But, however, when you look at what industry—what vertical—actually has, built-in, a large, a strong open-source culture and is contributing to open-source, like outside of tech—where it's roughly 50 percent of tech companies contribute to open-source, often on a regular basis, outside of tech I think the closest is finance and financial services, and it's like 12 percent, or like 13 or 14. It's really, really low. So, tech has it, the rest of the world, not yet.

And to some degree, that's also why open-source is actually a real accelerator of how companies are able to build the kind of tools that they need to respond to their business needs. It's not by accident that you find that the companies that have the highest growth—market growth as a company are those that are heavily invested in tech and heavily invested in open-source. And so it's not surprising that incumbents from all the verticals are having a much harder time to adapt, and as a result are also, in verticals where there's lots of competition, lots of new players, lots of new startups that are, sort of, like, stealing market share, and disrupting those different markets.

Emily: You've said a lot of things that are really interesting. I wanted to ask, though, again, about this idea of helping people develop soft skills because honestly, I had never considered that as an advantage of open-source. Could you just sort of talk a little bit about how that happens, and how individual engineers can use working on open-source projects to develop soft skills, and then how it translates to better success in their employment situation?

Tobie: So, if you look at how software is built in a closed-source project, you will essentially be working with your peers. I mean, that's not always the case, but in most cases, people that you can literally, like, turn your chair around and tap on the shoulder to get help. In open-source that's very different. Large open-source projects will have people across lots of time zones, and completely different stakeholders.

You will have in the same project, someone that is just passionate about this project and is a teenager in a high school that just really cares about whatever it is that you're working on. You’re going to have a bunch of folks in academia, actually using that project to run some data internally or something like that. You will have small companies building plugins on top of it or doing agency work. You will have large corporations leveraging that project. So, you will have this very broad stakeholder set of people with very different backgrounds, very different interests, very different reasons to be involved, essentially.

And I mean, just that, just this diversity of background and culture will make you up your communication game because you will not be able to speak to these very different stakeholders. If you want to get something out of them, if you want to review one of their pull requests, if you want to get them to sign the CLA, it's not the same as turning around and tapping your colleague on the shoulder that, unfortunately, tends to be roughly the same skin color, age, and gender as you are in lots of different teams, still today. So, I think that's the first point is just, lots of stakeholders, with lots of different interests, coming from lots of different places.

The second bit is, a lot of software is about communicating what you want to do and what you're hoping that they're doing. And that's harder to do in return, frankly, for most people. And it's harder to do, again, when you have sociocultural gaps. So, learning how to do that properly to get alignment on something, this is a skill, you have to learn. Thirdly, the absence of formal leadership in—which is what I was mentioning before—in projects and by formal leadership, I mean, yeah, sure, there's like a technical steering committee, or a [00:21:05 unintelligible], or someone's leading the project, like maintainers and stuff, but they don't get to tell who does what.

So, if you want help from someone on a project, you will have to learn how to use your soft skills to do that because you can't make anyone comply to anything. It's this completely soft, smushy thing. We don't really have—you can't hold on to someone and tell them, “Go do this PR now,” or, “Go review this.” You will have to figure out ways of getting people to be involved using a completely different skill set then force compliance.

And this is—I mean, I might be cutting corners here, but to me, this is what leadership is about. Leadership is about aligning people in the mission without a whip. And this is precisely what you're doing if you want to do anything in open-source. And this set of skills, once you’re back in a company—I mean, any kind of serious project, impactful project in a company will be across multiple teams, multiple orgs, you'll have to get approval from, like, policy, you'll have to go see legal, you’ll have to get designers involved, you’ll have to get product involved, you’ll have to get infrastructure invol—like, all of these organizations that you don't have direct power over, learning this set of skills inside of an open-source project prepares you for this so much.

Emily: Interesting. Yeah. You know, you could have a project, and legal could say, “No.” And you don't get to just override what legal says if they say no. You have to have the skills to negotiate a way out of that, basically.

Tobie: Yeah, and frankly, I mean, if you look at sort of the career ladder of an engineer, it's essentially around growing your impact inside of an organization or company. And growing your impact, I mean, that is done laterally. It's done by getting others aligned on your vision early in the process. And again, it's interesting because there's lots of parallel between what I'm describing right now and what I do for a lot of my clients, which is to get alignment from all of the different stakeholders along a specific set of goals. And this is only soft skills: it's listening to people, figuring out what their needs are—when legal says, “No,” I mean, no one ever says, “No.” People say, “No,” and they mean, “Oh, this is going to make me too—too big of a cost for me. I don't want to do this right now. It's easier for me to just say, ‘No.’ I don't really understand the risks. I don't really understand the value of this project.” I mean, behind the, “No,” there's a bunch of information. And building soft skills lets you have the tools to go figure out what that, “No,” that legal just gave you really is about. And it's way easier to address something like, “Oh, there's actually—I'm concerned about this specific risk at this specific place,” than addressing something that's as vague as, “No. Legal said no.”

Emily: And how would you say that an engineer that’s, say, in a non—technology company, as in you know, not in the technology vertical, they are in a company that sells cars, or pharmaceuticals, or financial services or whatever, what are the specific ways to to make those business arguments to talk effectively with business leadership?

Tobie: So, one of the consultants that is a consultant for consultants, David A. Fields, talks about ‘right side up thinking’ and he's essentially talking about, put yourself in the shoes of the people that you're talking about. Understand what it is that they care about, and then have answers to that. Which, to me, is also something that you can build in open-source, but it's essentially listening to people. I mean, there's so many times I've been in a meeting with lawyers about a particular topic for a client or for a company I was working with, where I got out of a meeting with much, much more than I expected, essentially because instead of opening my mouth, I just shut up and listened to what it is that they were concerned about, and really tried to understand from their perspective.

And then realize that all of the schemes I had in my head of what it is that they wanted, and the solutions I had for what I thought it was that they wanted, were not necessary at all; what they cared about was something completely different that I just couldn't know about. And so, that would be my biggest suggestion is, just shut up and listen to what people want. Same for customers, by the way. I mean, when you're facing customers, just actually listen to what people say.

It doesn't mean that you have to essentially implement precisely what solution they're giving you, but you have to listen to what their problem is. As an expert—and that's true of an engineer in a non-engineer context: the engineer is the expert, but your expertise should be applied to turn the need, the requirement, into something that's implementable. That's its only purpose, really. It shouldn't be about asking people, “Well, so would you want to use PHP or Rails for this?” And then giving them a lecture on both. This is not what someone some business wants to hear about.

Emily: Excellent. We are going to go ahead and wrap up pretty soon, but anything else that you would like to add about bridging the gap between business and engineering?

Tobie: Yeah, so I think that at the end of the day, what really works is when everyone is aligned, and pulling on the same rope, aligned with the same goals. If you're in a company, where the underlying goals really don't match at all your vision of what you want to do, you're in a bad place, regardless of what vertical that company is in, whether it is a tech or a non-tech company. So, I think that engineers, if they're able to and, again, I mean, not everyone is in position to change job or hop to find a different job, and the job market right now is particularly difficult, but I think that if you want to be happy in your job, you have to make sure that there's alignment. And if there isn't, at least try to carve out areas of alignment.

And don't try to win every fight: really go for the things that matter to you that make a difference, and make concessions. Actually, that's the other, for me, the really key point is, make concessions. If things don't really matter to you but make a huge difference for the person that's in front of you, make a concession even if you think it's silly. As engineers, we really have, again, this really binary way of thinking. Admit that there's a lot more to all of this then yes or no, and that there's a whole bunch of area in the middle where people can meet and find agreement, and focus on that stuff.

Emily: Excellent. All right, just a couple last questions. What is your favorite engineering tool that you couldn't live without?

Tobie: That's an interesting question. I don't think I really have one. I think that's deliberate. My goal would be to be able to jump on a new machine and be effective within seconds, and not have to go through the whole ordeal of having to set everything up just to right for me. So, I tend to try to work with whatever is there. I also don't believe that a good engineer is an engineer that types fast. Actually, I'm a really slow typer, so maybe that's why. But yeah, I really believe that it's not about tooling, it’s about all of the other things, and that tooling should come last. So, I don't have any is my answer.

Emily: Fabulous. And then where can we listeners connect with you or follow you.

Tobie: So, I tweet quite a bit under @tobie. So, T-O-B-I-E. There's lots of politics there, too, so if you believe that tech and politics are not linked, you probably don't want to follow my account. And then there's the website of my consultancy, which is unlockopen.com. So, unlock and open in one word, dot com.

Emily: Excellent. All right, well, thank you so much for joining me.

Tobie: Well, thank you for having me. This was fun. Thank you so much.

Emily: Thanks for listening. I hope you’ve learned just a little bit more about The Business of Cloud Native. If you’d like to connect with me or learn more about my positioning services, look me up on LinkedIn: I’m Emily Omier—that’s O-M-I-E-R—or visit my website which is emilyomier.com. Thank you, and until next time.

Announcer: This has been a HumblePod production. Stay humble.

View Details

The conversation covers:

  • Tracy’s thoughts on how the relationship between open-source and cloud-native should be described.
  • The advantages and disadvantages to an organization using open-source.
  • Some of the major risks associated with using open-source, and why companies should approach with caution.
  • Why CI/CD is a rising security concern for open-source organizations.Tracy also provides her thoughts on how businesses are handling the CI/CD pipeline today, and where the trend is heading.
  • Some of the unresolved challenges related to continuous delivery that currently exist.
  • Tracy’s advice for companies that are just starting to develop an open-source contribution strategy.
  • How companies should approach topics like open-source strategizing and building open-source communities.
  • The common mistakes that individuals and companies make when nurturing open-source communities. Tracy also comments on mistakes that people are making with continuous delivery.

Links

  • CloudBees: https://www.cloudbees.com/
  • Continuous Delivery Foundation: https://cd.foundation/
  • Twitter: https://twitter.com/tracymiranda

Emily: Hi everyone. I’m Emily Omier, your host, and my day job is helping companies position themselves in the cloud-native ecosystem so that their product’s value is obvious to end-users. I started this podcast because organizations embark on the cloud naive journey for business reasons, but in general, the industry doesn’t talk about them. Instead, we talk a lot about technical reasons. I’m hoping that with this podcast, we focus more on the business goals and business motivations that lead organizations to adopt cloud-native and Kubernetes. I hope you’ll join me.

Emily: Welcome to The Business of Cloud Native. Today, I'm chatting with Tracy Miranda. Tracy, thank you so much for joining me.

Tracy: Hi, Emily. Thanks for having me. It's my pleasure.

Emily: So, as usual, I just want to start off with having you introduce yourself, both what you do, where you work, but also, like, some details, what does this actually mean? How do you actually spend your day?

Tracy: Yeah, so I'm the director of open-source CloudBees, and I'm also the board chair at the Continuous Delivery Foundation, which is an open-source foundation, which is home to projects like Jenkins, and Spinnaker, and Tecton, and Jenkins X. So, basically, I'm a big fan of all things open-source, which in day-to-day means I'm doing anything which is related to building communities. So, either involved with code, or building communities and through conferences, or sometimes just the boring governance stuff around open-source.

Emily: What is the boring governance stuff around open-source?

Tracy: So, I guess it is just trying to get folks moving in the same direction, and reminding people that it's sometimes more than just code. And whether it's updating a code of conduct, and one of the things we've seen and—okay, I wouldn't call this boring; it's actually taken over a bit in open-source communities, but it's sort of different from the code, but it's the whole terminology updates. We've seen a lot of open-source communities have become more aware about wanting to be better about using terms like ‘master’ and ‘slave’ and move away from that. That being said, it's not that easy, so there's a lot to do in getting people on the same page and ready to move forward even before you can start changing a line of code.

Emily: Since the topic of the podcast is cloud-native, obviously, open-source and cloud-native are related. In fact, some people think that cloud-native must be open-source. Where do you fall on that spectrum? How do you think the relationship between open-source and cloud-native should be described?

Tracy: Yeah, I think that they're pretty distinct things. So, cloud-native is all about using the Cloud effectively and having technology which takes advantage of modern architectures to give you things like rapid elasticity, or on-demand self-service. And that's distinct from open-source, which is around the licensing, and it's become more about communities, as well. But I think because Kubernetes has been the most successful cloud-native project that is open-source, I guess there's become this very, very strong association which, in my mind, is a very, very good thing because I think open-source communities are really the way to drive innovation very, very quickly across the industry.

Emily: And this may seem sort of obvious, but what are some of the advantages and disadvantages to an organization in using open-source?

Tracy: Yes. So, I think—well, lots—virtually every company uses open-source, and the first thing people can see as the benefits are just the engineering efficiencies. So, using technologies which, say aren’t core to the business, but then building on top of those and taking advantage of the features rather than dedicating their own engineering resources to developing them. I used to work as a consultant, and I would go from company to company, and usually, they would be adopting open-source when they wanted to get away from an in-house project where the people or person who had written it had left the company. So, I think there's a lot to be said, as well, for sustainability of technology: that communities and open-source communities are really good at sustaining projects over the long term, and therefore kind of the best bet for technology that's going to live on beyond individuals or even companies, acquisitions, or whatever.

Emily: Do you think there are any risks to using open-source? I'm even interested in hearing if there are risks that are not real, but that are perceived risks. And then even maybe some risks that people don't think about, but that are in fact, quite real.

Tracy: Yes, yeah, no, absolutely there are risks. So, it's wise for companies to approach with caution. I think the risks sort of depend on which side—like, are you looking to just use open-source that someone else has written, or are you contributing something, which might be key to your company, but then you’re saying, “Okay, I'm going to do this in an open way,” which brings us to one of those common perceived myths, that someone, like a cloud provider, is then going to take your open-source software and do a better job of making money around it, so thereby just ruining your entire business model.

And I think the other area where we tend to see a lot of dialogue around, is always around open-source security. For a long time, people used to, sort of, make out that this was different from closed source security, somehow. Security through obscurity meant that closed-source was better than open-source, which is clearly not the case. You can have secure open-source software, not secure open-source software. It just really depends on the project and the practices.

Emily: And then also, I thought we'd talk a little bit specifically about this CI/CD work that you do. How important is CI/CD, do you think, in the pursuit of being cloud-native?

Tracy: Yes, no, I think CI/CD has just risen to the top as one of the key concerns. And I think, part of the reason—when you're doing things in a cloud-native way it means that your systems are very distributed; you don't necessarily know where the services are running, it's typically not on-premise, and suddenly it becomes very important to understand how do you do this integration, and how do you then deliver that software in a way that is both quick, and that is not going to—you can do it in a safe way, so it's not going to break every time you do releases. And I think we're seeing that it really is at the forefront. Like last year, we started the Continuous Delivery Foundation, which is an open-source foundation, and the mission there is to increase the world's capacity to ship software securely and at speed. And the uptake from folks has been really well. Everyone's grappling and trying to figure out, what does CI/CD look like in the Cloud? What does it mean to be cloud-native CI/CD?

Emily: And from the perspective of an end-user, what do you think are some of the, still, unresolved challenges related to continuous delivery?

Tracy: Yeah, it's very challenging. Everything is changing under enterprise’s feet. And it's not just the tools we're using, is also the skills we expect people to have, the way we organize a team. And traditionally, it's been very, very hard to decommission software or deprecate it, but what we're seeing in the industry now is that everything is changing really rapidly. You take something like Kubernetes and it has a new release, like, every three months and then nine months later, that's deprecated.

So, people are having to make changes in enterprise situations at a rate that they just previously didn't come anywhere close to, and that's pretty challenging when you're having to deal with the changing tools, and processes, and people all at the same time, all while keeping your business up and running.

Emily: In terms of the whole CI/CD pipeline, do you think most end-users experience that as being mature? Is it sort of figured out, or is it something that they continue to struggle with?

Tracy: I think everybody has a CI… certainly CI… many people have sort of cracked, and they've got their systems set up. And then the delivery side, it just, kind of, varies. And I think it depends; we see a lot of folks who are really trying to figure out pipelines and are really trying to figure out what that looks like in a cloud-native world, and they haven't figured out, what does it mean for things to be highly available? What does it mean to be able to scale at any level? So, everybody's got something, but I think we've only just scratched the surface of what's possible with today's technology.

Emily: Where do you think it's going in the future?

Tracy: Yeah, I think, like in the same way we're having this big shift, everybody's got monoliths, and the problem with the monolith is that you can't do the speed and security at the same time. So, if you think about the key metrics people use today, there’s two on speed, “Which is how quickly can you deploy?” And, “What's your lead time for changes?” And then for the safety, it's, “How long would it take you to restore services, if something went wrong?” And, “What is your change failure rate? How often are things going wrong every time you push code?”

So, in the bid to get really good at those metrics, I think people have realized that monoliths cause a lot of problems, and it's much easier to meet these capabilities if you've got microservices are smaller batches of code, each, which do a specific thing, and there's less chance of things falling over when you make changes because there's not all these huge dependencies. Now, however, when you do start having all these different microservices with, let's say, a web of dependencies, things start to get really complicated. So, now you don't have, perhaps, one CI/CD pipeline, you have a pipeline per microservice. And then we start to say, “Okay, what is the definition of the application even? Is it all these microservices? Which version is it?”

And then things like configuration management start to enter the picture, especially if you've got dependencies on things, let's say, outside your company, or open-source. So, I think it's a lot for people to grapple with, like, how to truly do microservices, and how the definition of an application is going to evolve. And I think for CI/CD, we can't keep doing what we've done in the sense of traditionally, folks have written a pipeline by hand, and you'd write a pipeline for your monolith. But now you've got all these different microservices. You want to start thinking about how can you have a pipeline auto-generated for them.

Emily: I wanted to actually shift and talk more about open-source communities as well since I know that's a large part of what you do. My first question is, what would you say to a company that's starting to think not just about consuming open-source, but developing a strategy to contribute to open-source? What do you advise companies who are just starting that journey?

Tracy: Yeah, no, I think for companies, it's a really good thing. I think open-source can give you a lot of strategic advantages, especially if you're coming in strong, and you're looking to be a leader in a space. And if we talk about category creation, you can use open-source almost as a weapon to drive the industry in a specific direction. So, I think what is important for companies is to be very deliberate about this strategy because open-source strategies can be almost counterintuitive, especially to folks who haven't done it before. This idea that you're giving away assets for free, or making them open. So, it's really important to have all the stakeholders in the company on the same page, and really understanding that this is a long-term thing where you'll have these benefits and not something where you start off and you do sort of half-heartedly.

Emily: Are there two or three, sort of, primary open-source strategies?

Tracy: Yeah, no, I think—[00:13:42 unintelligible] I think you can break it down. So, people would talk about the Red Hat model, which is really hard to reproduce but everything was open-source, and then they have this whole—they layered on top of that with a lot of services, and things. And then there's the open-core model where you're separating an open-source portion of the product, and then you add on a lot of features and things that add value that aren't being produced in the open-source. So, I think there's those, and then the new one that we're starting to see more of is—just looking much more at SaaS platforms. So, you have some open-source code, but your real—where you're making money is by offering it as a service.

Emily: And how does that differ for a company whose core business isn't software? So, for example, if you're something like a Home Depot, and almost undoubtedly you use open-source software. If Home Depot wants to start contributing as well, as part of their company strategy, what should they know? What should a company like that think about as strategies?

Tracy: Yeah, no, I think that's a great point because we do see a lot of companies contributing, and actually a lot of innovation is coming from companies who use software, but they have a different focus. And I think one good example, as well, is Capital One, who have a lot of open-source they contribute and maintain. And it's different, it's separate from, kind of, the main banking function. So, I think, again, for companies like that, it's just mapping out the strategy, being very deliberate in is there some sort of monetization around this, or is it more—you know, we see a lot of companies who want to do it to be seen as leaders in the field, and to, sort of, share some innovation to be seen as an attractive place, as well, for people to work with, and just to really drive that industry to help the innovation and to help make it a good place to be. So, I think the same things apply there, although maybe the business models allow, perhaps, for a bit more freedom. And we often find in those companies, they will have open-source program offices, which is a dedicated set of people who will map out the strategy and pull the whole company along in the same direction.

Emily: Obviously, a big part of open-source is building a community. How do you do this? How do you herd the cats in a way that advances your project? And I'm actually curious, I don't know if you have a perspective on this from both somebody—an individual starting a project, and a company that wants to create a community around a particular project?

Tracy: Yeah, no, I think that's a really great question. And people are always attracted to, I think, you want to start out with the big idea: why is your project going to do things better than before, or what's nicer about it? So, I think you have to start with, I guess you'd call it, like, you're [00:16:58 unintelligible] for your open-source project; the reason people are going to be attracted to it, and they're going to come and say, “Actually, I want to be part of this.” Because I think people do want to feel part of something bigger than themselves. They also want to see other people contributing, and everybody pulling their weight, and not necessarily any kind of biases for specific companies.

So, the more open you can make it, the more transparent you can be about how things happen, people love to—if they're committing, and folks in open-source do commit fully—they want to know that they're not going to be taken advantage of, that they can do that, and they can really change the way the project is going to—they can feel the change they're going to make. So, I think it's important just to go to those principles of openness and transparency, and to let people participate. I think sometimes having clear ways—like with Jenkins, we saw that originally it really thrived because people could write their plugins, and they could make it their own, and they could share them and show them to their friends. And it's the same idea with GitHub, things that make developers look good as well, while they're contributing to open-source always makes for very, very successful projects.

Emily: What do you think are common mistakes that people—individuals or companies make around nurturing the community?

Tracy: Yeah, I think the mistakes are always connected to control and wanting to control too much or in a too specific way. And you could almost—I don’t know if this is a good analogy, but it's almost like, I guess, parenting, in a way. You might be tempted to be very regimented and say, “Okay, your child can do this, or they can't do that.” But then you sort of lose out in finding out where could this go? How big could this grow?

So, I think it's finding the right level of control so that the project can take on a life of its own and be used in ways that perhaps you couldn't even imagine. I think that's when the real magic happens. But it does take a leap of faith and understanding that you will be able to reap some business benefit out of this if that is your aim as well.

Emily: Do you think that that's easier for individuals or for companies to achieve?

Tracy: I think it depends on what people are going into it for. And for individuals, I think often it's they want to share their idea with the world or they want to build a reputation, which is very synonymous with doing the project. Having said that, individuals can have the same issues around wanting to control it, but I think there's perhaps a different monetization emphasis which would make it easier.

Emily: Actually, I had a similar question related to continuous delivery which is, do you find that there are common mistakes that you see people making?

Tracy: Yes. And some of the mistakes, I guess—one of the most common mistakes is a pretty boring one. And I know why it happens, but [laughs] it's just around documentation, to be honest. And it's the, “Okay, we're going to write the code, and then we're not going to necessarily document it or share the way people can either get involved or use a project.”

And it's just—documentation is hard. Good documentation is really hard. Things keep changing, and it's boring to go keep updating them. But it's so incredibly important, and some of the most successful open-source projects have always provided that kind of self-service set of docs where people don't have to be asking the same questions over and over again. They really can go off and feel empowered to do things and to do things and not feel like they're getting it wrong or wasting their time, which I think is really important when building community. So, yeah, just write good docs, everybody.

Emily: And do you think there's anything else specifically related to how companies approach continuous delivery, that there's something that a lot of them are not doing right?

Tracy: With continuous delivery, especially today where everybody's in a really, kind of, tricky situation where they're trying to make this move to using cloud-native technologies because the benefits are so huge, but at the same time, all these technologies are coming very thick and fast, and nobody's sure—people have tried technologies which are now no longer used, so this is a bit of fear of saying, “Okay, is this going to be a safe bet? And at the same time, while I'm trying to decide if that's the right technology to use, I'm having to restructure my teams, and change of habits is really hard, and we've got all these additional environments we're having to deliver software for.” So, it's a huge challenge, and everything has to be done in balance: you have to get the tools, and you have to get the technology, and you have to get the people right. You can't do any one of those and hope it's going to work, you have to do this juggling act within your organization. And that's massively, massively challenging, especially when you are trying to change long-held behaviors and habits people have, and just ask them to do things in a different way.

Emily: Do you think technology is more challenging, or people skills organization is more challenging?

Tracy: Yeah, I think the thing with technology that is more challenging today is, especially in the CI/CD space, we have a lot of different types of tools. And we don't have standard ways to talk about—like, we don't have standardization of terms, so different things have different meanings to different people. So, you might say ‘a pipeline’ but it might mean—the scope might change depending on who you're talking to. And so it's really hard for people to understand, how do I connect these different tools together? There's very poor interoperability, as well, which is another thing the Continuous Delivery Foundation wants to try and solve.

So, I think those are key areas. Security is another one, which makes it really hard when you break things up. And no one's taking responsibility for the interaction between different platforms of different open-source technology written by different people, that becomes really tricky. So, I think we do need solutions at a community level, and we need communities working together closer to tackle this proliferation, and lack of interoperability, and new security concerns that we have to deal with as an industry.

Emily: Is there anything else that I didn't think to ask that you'd like to add?

Tracy: Yeah, no. I think what we're doing in the Continuous Delivery Foundation, if I can say a little bit about that, it is a relatively new open-source foundation. And I think it's a good place to bring people together where we're trying to tackle these issues. So, things like interoperability, we have an interoperability working group. And one of the first things that happened in that group as people would come together and talk about the different tools, is that we spontaneously realized we needed to define the tools.

And there was a page set up where everybody could write down the definition of how their tool—use different terms. You know, is it a step? Or what do you call it in your tool? So, we have this what we call, like, the Rosetta Stone, of CI/CD tools. So, it compares across—whether it's all kinds of Git providers or pipeline orchestration tools, was the different terminology.

And I think from there, we're going to look to see how we can standardize as an industry, just to make it simpler for people because I think—I would really hate to be someone new coming into the industry today and trying to figure out where to start, which tool to try out because the amount of noise and confusion is at all-time high levels.

Emily: That's absolutely fair. And in fact, speaking of tools, my next question is, what tool do you really rely on? What engineering tool would you not be able to work without?

Tracy: Yeah, well, they kind of say for developers, and I think this rings true for me as well, you’re kind of in three places. You’re in, like, GitHub and Slack, and then your development environment which use VS code, and like many people. So, those are, kind of, the three development environments. I think, when I look at CI/CD, and we look at new technology in the space that’s, kind of, gaining quick adoption, there's two projects in CDF which are starting to really resonate. And one is Tekton, which came out of Google, and their Knative serverless platform.

But that's looking to have these standardized building blocks for CI/CD pipelines. And then the other one is Jenkins X, which, incidentally, uses the building blocks of Tekton to stitch together a CI/CD experience, if you wish, that pulls in Kubernetes, and Helm, and all those other projects to give a really nice developer experience just generating pipelines for you, so you don't have to write things by hand, and giving you preview environments, and really just trying to take advantage of all the power that cloud-native affords you in delivering software.

Emily: And then lastly, how can listeners connect with you or follow you?

Tracy: Yeah, no, I think the best place is Twitter. So, find me Twitter at @tracymiranda, and in all the continuous delivery working groups, and the communities we're building there. So, find that on cd.foundation, and, yeah, come join the community. We're having some great conversations in the space.

Emily: Well, thank you so much, Tracy, for joining us.

Tracy: Yeah, thanks for having me. And yeah, really great conversation and questions.

Emily: Thanks for listening. I hope you’ve learned just a little bit more about The Business of Cloud Native. If you’d like to connect with me or learn more about my positioning services, look me up on LinkedIn: I’m Emily Omier—that’s O-M-I-E-R—or visit my website which is emilyomier.com. Thank you, and until next time.

Announcer: This has been a HumblePod production. Stay humble.

View Details

The conversation covers:

  • The main function of an OSPO, and why Comcast has one.
  • How Nithya approaches non-technical stakeholders about open-source.
  • Where the OSPO typically sits in the organizational hierarchy.
  • The risk of ignoring open-source, or ignoring the way that open-source is consumed in an organization.
  • Why every enterprise today is using open-source in some way or another.
  • The relationship between cloud-native and open-source.
  • Some of the major misconceptions about the role of open-source in major companies.
  • Common mistakes that companies make when setting up OSPOs.
  • Why Nithya and her team rely heavily on the TODO Group in the Linux Foundation.

Links:

  • Comcast: https://www.xfinity.com/
  • Linux Foundation: https://www.linuxfoundation.org/
  • TODO Group and The New Stack survey: https://thenewstack.io/survey-open-source-programs-are-a-best-practice-among-large-companies/
  • Trixter GitHub: https://github.com/tricksterproxy/trickster
  • Kuberhealthy GitHub: https://github.com/Comcast/kuberhealthy
  • Comcast GitHub: https://comcast.github.io/
  • Nithya Ruff Twitter: https://twitter.com/nithyaruff

Transcript

Emily: Hi everyone. I’m Emily Omier, your host, and my day job is helping companies position themselves in the cloud-native ecosystem so that their product’s value is obvious to end-users. I started this podcast because organizations embark on the cloud naive journey for business reasons, but in general, the industry doesn’t talk about them. Instead, we talk a lot about technical reasons. I’m hoping that with this podcast, we focus more on the business goals and business motivations that lead organizations to adopt cloud-native and Kubernetes. I hope you’ll join me.

Emily: Welcome to The Business of Cloud Native, my name is Emily Omier, and today I'm chatting with Nithya Ruff, and she's joining us from the open source program office at Comcast. Nethya, thank you so much for joining us.

Nithya: Oh, it's such a pleasure to be here, Emily. Thank you for inviting me.

Emily: I want to start with having you introduce yourself, you run an open source program office. And if you could talk a little bit about what that is, and what you do every day.

Nithya: So, just to introduce myself, I started working in open-source back in 1998, when open-source was still kind of new to companies and organizations. And from that point on, I’ve been working to build bridges between companies using open-source and communities where open-source is created. At Comcast, I have the pleasure of running our open source program office for the company, and I also sit on the board of the Linux Foundation and recently was elected chair. So, it gives me a chance to both look at the community side through the LF and through corporate use of open-source at Comcast.

So, you also ask what does an OSPO do? What is an OSPO, and why does Comcast have one? So, an open source program office is a fairly new construct, and it started about 10, 11 years ago, when companies were doing so much open-source that they really couldn't keep track of all of the different areas of open-source usage, contribution, collaboration across their companies. And they felt that they wanted to have a little more coordination, if you will, across all of their developers in terms of policy for use, the process for contribution, and some guidelines around how to comply with open-source licenses and, on a more strategic note, to educate both executives as well as the company in terms of open-source and opportunities from a business engagement and a strategy perspective. So, you find that a lot of large companies typically have open source program offices.

And we, frankly, have been using open-source for a very long time as a company, almost since the turn of the century, around 2005. And we started contributing and our number of developers started growing, and we didn't realize that we needed a center of excellence, which is what an open source program office is, where people can come to ask for help on legal matters—meaning compliance and license matters—ask for help in engaging with open-source communities, and generally come for all things open-source; be kind of a concierge service for all things open-source.

Emily: And how long has Comcast had an OSPO?

Nithya: I came on board in 2017 to start the OSPO, but as I mentioned before, we’ve done open-source organically throughout the company for many, many more years before I came on board. My coming on board just, kind of, formalized, if you will, the face of open-source work for the company to the outside world.

Emily: You know, when we think about open-source in the enterprise, what sort of business opportunities and risks do you have to balance?

Nithya: That's a great question. There are lots and lots of great business value and opportunity that companies get from open-source. And the more engaged you are with open-source, the more business value you'll get. So, if you're just consuming open-source, then clearly it reduces the cost of your development, it helps you get to market faster, you're using tried and tested projects that other companies have used and hundreds of developers around the world have used. So, you get a chance to really cut cost and go to market faster.

But as you become more sophisticated in collaborating with other companies and contributing open-source back, you start realizing the benefit of, say leveraging a lot of other developers in maintaining code that you've contributed. You may start off at contributing a project, and you are often the only one bearing the burden of that project, and very soon, as it becomes useful to more and more people, you're sharing the burden with others, and you benefit from hundreds of new use cases coming into the code, hundreds of new features and functions coming in which you could never have thought of as a small team yourself. I believe that the quality of code improves when you're going to open-source something, it helps with recruitment and thought leadership because now candidates can actually see the kind of work that you do and the quality of work that you produce, and before that, they would just know that you were in this space, or telecom, or other areas, but they could not see the type of work that you did. And so, to me, from a business value, there's a tremendous amount of business value that companies get.

On the risk side is the fact that you need to use it correctly, meaning you need to understand the license; you need to understand how you're combining your code with the proprietary code in your company; you need to understand if the code is coming from a good community, meaning a healthy community that is here to stay, and that has a good cadence of releases and is vibrant from an activity perspective; you need to also understand that you need to be engaged with the open-source community and understand where that particular project is going and to be able to sit at the table to influence or contribute to the positive direction of that project, and sustainability of that project. So, if you just consume and don't engage, or don't understand the license implications or contribute, I think you're not getting all of the value and you risk being considered a poor citizen in the community. And frankly, if people don't collaborate with you or cooperate with you, sending a patch upstream may take months to be accepted, as opposed to someone who's part of the community, who's accepted, who's seen as a good citizen. So, I think you've got to invest correctly through either an open source program office or a really intentional and thoughtful program to engage with the community in order to really mitigate risk, but also get the full benefit of working with open-source.

Emily: And what do you find you have to educate the non-engineering stakeholders about, so the business leadership when you're talking about open-source?

Nithya: That's also a very important function of an OSPO in my mind, is really making sure you have executive sponsorship and business buy-in for why open-source is a key part of the innovation process in the company. Because as you correctly said, there is a level of investment one needs to make, whether it is in an OSPO or in the compliance function, or for engineers to take the time to upstream their patches or to engage with communities. It all takes investment of time and money. And business needs to buy into why this is a benefit for the company, why this is a benefit for the business. And very often, I find that leadership gets it.

In fact, some of my best sponsors and champions are executives. Our CTO, Matt Zelesko, completely gets why open-source is important for the business innovation, competitive advantage. And so, also my boss Jon Moore gets it. And I found that in a previous company where I was starting an open source program office, I had to work a little harder because it was a hardware company and they did not understand how working in open-source would fit into the engineering priorities. And so we had to, kind of, share more about how it allowed us to optimize software for our hardware, how it allowed us to influence certain key dependencies that we had in our product process, and that customers were asking for more open-source based software on our product.

So, yes, building the business case is extremely important, and having sponsors at the business is extremely important. The other key constituent is legal, and working with your legal team hand-in-hand, and understanding their role in assessing risk and sharing risk with you, and your role as a business saying, “Is this an acceptable risk that I want to take on? And how do I work with this risk, but still get the benefits?” is as important. We have a great legal team here, and they work very closely with us, so we act as the first line of questions for our developers.

And should they have any questions about, “Should I use this license? Or should I combine this license with this?” we then try to give them as many answers as we can, and then we escalate it to legal to bring them into the discussion as well. So, we act as a liaison between us and legal. So, to your point, it is important for the business to understand. And the OSPO does a great job in many of my companies that I’ve worked with to educate and keep business informed of what's happening on the open-source side.

Emily: You mentioned working with developers; what is the OSPO's relationship with the actual developers on the team?

Nithya: So, we don't have many developers on our team. In my OSPO, I have one developer who helps us with the automation and functioning of the open source program office tools and processes. Most of us are program managers, and community managers, and developer relations managers. The developers are our customer. So, I think of the developers in Comcast as our customer, and that we are advocates for them.

And their need to use open-source in a frictionless way in their development process as our objective. So, we've worked very, very hard to make sure that the information they need, the processes they need are well oiled, and that they can focus on their core priority, which is getting products to market to really help our customers. And they don't need to become experts at compliance, they don't need to become experts at any of the functions that we do. We see them as our customer, so we act as advocates.

Emily: Where exactly in the organizational hierarchy, or structure is the OSPO? Is it part of the engineering team?

Nithya: Yes. I think that is the best place for an OSPO reside because you really are living with the engineering organization, and you’re understanding their pain points, and you’re understanding the struggles that they have, and what they need to accomplish, and their deadlines, et cetera. So, we live in the product and technology organization, under our CTO, who's also part of the engineering organization. So, I find that the best OSPOs typically reside in engineering or the CTO office, there are some that reside in legal or marketing. And whenever you decide, it tends to flavor the focus of your work. For us, the focus of our work is how can we help our developers be the best developers and use the best open-source components and techniques to get their work done?

Emily: And what do you see is the risk that organizations take if they ignore open-source, they don't have this, sort of, conscientious investment in either an OSPO or some other way to manage the way open-source is consumed?

Nithya: This is how I would put it. Everything that you see in technology development today, a lot of the software that we consume, whether it’s from vendors, or through the Cloud—public cloud, private cloud—is made up of open-source software. There’s a ton—I would say, almost 50, 60 percent of infrastructure software, especially data center, cloud, et cetera, is often open-source software. So, if you don't know the dependencies you have, if you don't know the stack that you're using and what components you have, you're working blindly. And you don't know if one of those stack’s components is going to go away or going to change direction.

So, you really need to be cognizant of knowing what you're using, and what your dependencies are and making sure that you're working with those open-source communities to stay on top of your dependencies. You're also missing out on really collaborating with other companies to solve common problems, solve them more effectively, more collaboratively. It's a competitive advantage, frankly, and if you don't intentionally implement some sort of an OSPO, or at least someone is tagged with directing OSPO type of work in the company, you're missing out on getting the best benefits of open-source.

Emily: Do you think there are any enterprises that don't use open-source?

Nithya: No. I believe that every single enterprise, knowingly or unknowingly, have some amount of open-source in their product, or in their tools, or in their infrastructure somewhere.

Emily: And what percentage of enterprises have—this is obviously just going to be your best guess, but what percentage of enterprises have an OSPO?

Nithya: I think it's a small percentage. New Stack and the TODO Group do a very, very good survey. I would refer us to that survey. And that gives you a sense of how many companies have an OSPO. I believe it's something like 45, 50 percent have OSPOs, and then another 10, 15 percent, say we intend to start one in the next two to three years. And then there's another, I don't know 30 percent that say, I have no intention of starting one. And the reason may be because they have a group of volunteers or part-time people across their organization who are fulfilling those functions between their legal team, and a couple of expert developers, and their communications team, they may think that they have solved the problem, so they don't need to have a specialized function to do this.

Emily: I wanted to ask a little bit about the relationship between cloud-native and open-source. What do you see as that relationship?

Nithya: If you ask anyone—and this is my opinion as well—that cloud-native technologies are very open-source-based. Look at Kubernetes, or Prometheus, or any of the technologies under the CNCF umbrella, or under any of the cloud-native areas, you find that most of them have their roots, or are created in the open-source way of development. So, it is an integral part of participation in cloud-native is knowing how to collaborate in an open-source way. So, it makes a lot of sense that CNCF is under the Linux Foundation, and it operates like an open-source project with governance, and technical body, and contributors. So, for us as well, being a cloud company—or a company that uses Cloud to host our infrastructure, and also a user of public cloud, we think that knowledge of open-source and how to work with open-source helps us work more effectively with the cloud ecosystem.

And we have contributed components like Trickster, which is a Prometheus dashboard acceleration component. We've also contributed something called Kuberhealthy, which allows you to really orchestrate across Kubernetes clusters to open-source because we know that that's the way to function, and influence, and if you will, kind of take advantage of the ecosystems in the cloud-native technology stack. So, cloud-native is all built on open-source. So, that's the relationship in my mind.

Emily: Yeah. I mean, I think actually, the Linux Foundation defines cloud-native as built on open-source software. I forget the exact words.

Nithya: Yep. I think so, too.

Emily: What do you think are some misconceptions out there, particularly among the enterprise users, about open-source and about the role of open-source in a major company?

Nithya: There are a number of misconceptions. And we talked and touched upon a few before, but I think it's worth repeating it because you need to confront these misconceptions and start engaging with open-source if you want to compete with the other companies in your industry, who all are becoming digital companies and are digitally transformed. And they need to work with open-source as part of their digital framework. So, one of the misconceptions is that vendor-supplied software or products don't have any open-source in them.

In fact, a lot of vendor-supplied software, maybe even from Microsoft has some open-source in them. Even from Apple, for example. If you look at the disclosure notices, you'll see that all of them consume open-source. So, whether you like it or not, there is open-source and you need to understand and manage it.

The second is not knowing what your engineers are downloading and using, and hence what you're dependent upon as a company, and whether those components are healthy, and whether those communities are doing the right thing. You need to understand what you're using. It's like a chef: you need to know your components, and the quality of the food that you create will depend upon the components you use. You'll also need to understand licenses and watch needed to comply with those licenses, and need to put process in place to comply with those licenses.

You also need to give back; it's not enough to just consume and not contribute back things like bug fixes, patches, and changes you make because you end up carrying all of that load with you as technical debt if you don't upstream it. And, frankly, you also consume, so you should give back as well. It's not sufficient to just take but not give. The last one is that open-source is free.

And so, many people are attracted to open-source because they think, “Ah. I don't have to pay any license fees. I can just get it, I can run it anywhere I want, and I can change it,” et cetera. But the fact of the matter is if you want to use it correctly, you do need to invest in a team that knows how to support itself, knows how to work with the community to get patches or make change happen, you need to build that knowledge in the house, and you do need to have some cost of ownership associated with using open-source. So, these are some of the major misconceptions that I see in companies that are not engaging with open-source.

Emily: And what do you see as, in your experience, some of the mistakes that companies can make, even when they're in the process of setting up an OSPO? What have you learned—maybe what mistakes have you made that you wish you could go back in time and undo—and what advice would you give to somebody who was thinking about setting up an OSPO?

Nithya: Couple of mistakes that come to mind is releasing a piece of code that's not been well thought through, or properly documented, or with the correct license. And you find that you get a lot of criticism for poor quality code, or poorly released projects. You end up not having anyone wanting to work with the project or contributing to the project, so the very intent of getting it out there so that others could use and collaborate with you is lost. And then sometimes companies have also made announcements saying that they want to release a particular piece of software, and they backtrack and they change their mind and they say, “No, we're not going to release it anymore.” And that looks really poor in the community because there are people who are depending upon it or wanting it, and it can affect the reputation of a company.

There was one more thing which I was going to say is, is really not being a good player. For instance, keeping a lot of the conversations inside the company, in terms of governing a project or roadmap for a project, and not being transparent and sharing the direction of the project or where it's going with the community. For an open-source project, is really bad. It can affect how you're perceived and how you're trusted or not trusted in the community. So, it's important to understand the norms of open-source, which is transparency, collaboration, contribute small pieces often, versus dumping a big piece of code or surprising the community.

So, all of these things are important to consider. And, frankly, an OSPO, helps you really understand how community behaves: we often do a lot of education on how to work with community inside the company, and we also represent the company's interests in communities and foundations and say, “This is where we are going. This is where we need your help.” And the more transparent you are, the better you can work with community. So, those are some areas where I've seen companies go wrong.

Emily: And when a developer who works at Comcast contributes to a project is he or she contributing as an individual or as part of the company? And how is that, sort of, almost, tension navigated?

Nithya: Most companies have a policy that any work that you do during your workday or on work equipment, is company property, right? And so it's copyrighted as Comcast, and most of our developers will contribute things under their Comcast email id. And that's fairly normal in the industry. And there are times when developers want to do work on their own time for their own pet projects, and they can do it under their personal emails and their personal equipment. So, that's where the industry draws the line.

Of course, there are some companies that are very loose about this type of demarcation, and some companies are incredibly tight depending upon the industry they're in, regulated versus high tech. But we are very encouraging of our developers to contribute code, whether on their own time or during company time, and we make the process extremely easy. We have a very lightweight process where they submit a request to contribute, and the OSPO shepherds that contribution through legal, through security, through other technical reviewers, and all in the interest of making sure we provide guardrails for the developer so that he or she does it successfully and looks good when they make the contribution. So, 95 percent of the time, we approve requests for contributions. So, very, very rarely do we say, “This is not approved,” because we think it's the right thing to do to give back and to share some of the work that we do with others, just like we get the benefit of using others’ work.

Emily: Is there anything else that you want to add about what OSPOs do, what they bring to the business, the relationship between cloud-native and open-source, anything that I haven't thought to ask that I should have?

Nithya: The OSPO, if you will, is a horizontal function that cuts across the entire enterprise development and helps coordinate and direct the intelligent and judicious use of open-source. So, that's why it touches all of the software development tools, apps, vendor-supplied software, public clouds, internal clouds, et cetera. Wherever open-source is used, which is everywhere, we touch it. And we also serve as the external face of the company to the open-source community so that the open-source community has one place that they can come to for questions, or to give feedback on something they're doing, or to ask a license question, or to ask for sponsorship or support for a conference or a foundation. So, it really makes open-source navigation very, very effective for the community, as well as for inside the company.

So, I'm a huge, huge fan of OSPO. I also love running an OSPO, and the kind of people that are typically in an OSPO. They tend to be very versatile, very general, they can pivot from legal to development matters to marketing and communications to really assisting a developer navigate something challenging. So, they're very versatile and terrific type of people. They also tend to have very high EQ and tend to make sure that they have a service mentality when they take care of questions that come in. So, I would say an OSPO is a great role for someone who wants to help and wants to know the breadth of software development.

Emily: It sounds like you're making a recruitment pitch.

Nithya: Uh, yeah. I don't have any openings right now, but I'm always encouraging and mentoring other OSPOs. I do at least one or two consultations with other OSPOs because we enjoy what we do as an OSPO and we want to help other OSPOs be successful.

Emily: I mean, is it hard to find people to work in OSPOs?

Nithya: It's kind of hard, in the sense that there are not too many people who do this work. So, I know, practically, I know all of the OSPO leadership and people who do this line of work in the industry. And it takes—some who come from a developer background. They have grown up as a developer using open-source and know the pains that they faced inside their company using open-source and not having certain processes or certain support or tools, and they go out to change the world in that way.

I came from a different direction. I came from strategy and product management, and I came with the notion of, “How do I connect the dots better across the organization? How do I make sure that people know what to do and how to build relationships?” So, I came from that perspective. Frankly, I think it's something that innately people have, which is the ability to absorb a lot of different types of knowledge and connect the dots and work to change things. You don't have to be born in open-source to be a good OSPO person. You just need to have a desire to help developers.

Emily: Was there any tools—it doesn’t, obviously, have to be a software development tool, but any tools that you could not do your work without?

Nithya: More than tools, I would say the organization that we rely on very heavily is the TODO Group in the Linux Foundation because it is a group of other OSPO people. And so it's been a great exchange of ideas and support, and tips, and best practices. The couple of tools that we use very, very heavily, and the love using, clearly, is something like GitHub or GitLab which helps you coordinate and collaborate on software development and documentation, et cetera. The other tool we use a lot to build community inside the company is things like Slack, or Slack equivalents because it helps you create communities of interest. So, when we are doing something around CNCF, we have a CNCF channel. We have a very, very large open-source channel that people come in and ask questions, and the whole community gets involved in helping them. So, I would say those are two really good tools that I like, and we use a lot in our function. And the TODO Group I think is a fabulous organization.

Emily: And where should listeners go to learn more and/or to follow or connect with you?

Nithya: There are two places I would say. comcast.github.io is where we publish all of our open-source projects, and you can see the statement we make about open-source. We also feature job openings at Comcast as well as our Innovation Fund, which is a grant-based fund request, so people can make a request for us to contribute money towards their project, or to research. And I'm on Twitter at @nithyaruff.

Emily: Well, thank you so much, Nithya. This has been really fabulous.

Nithya: Thanks, Emily. And thank you for helping me share my enthusiasm for what an open-source office is, and why everybody needs one.

Emily: Thanks for listening. I hope you’ve learned just a little bit more about The Business of Cloud Native. If you’d like to connect with me or learn more about my positioning services, look me up on LinkedIn: I’m Emily Omier—that’s O-M-I-E-R—or visit my website which is emilyomier.com. Thank you, and until next time.

Announcer: This has been a HumblePod production. Stay humble.

View Details

This conversation covers:

  • The advantages of using a distributed data storage model.
  • How Storj is creating new revenue models for open-source projects, and how the open-source community is responding.
  • The business and engineering reasons why users decide to opt for cloud-native, according to Ben.
  • Viewing cloud-native as a journey, instead of a destination — and some of the top mistakes that people tend to make on the journey. Ben also talks about the top pitfalls people make with storage and management.
  • Why businesses are often caught off guard with high storage costs, and how Storj is working to make it easier for customers.
  • Avoiding vendor lock-in with storage.
  • Advice for people who are just getting started on their cloud journey.
  • The person who should be responsible for making a cloud journey successful.

Links:

  • Storj Labs: https://storj.io/
  • Twitter: https://twitter.com/golubbe
  • GitHub: https://github.com/golubbe

Transcript

Emily: Hi everyone. I’m Emily Omier, your host, and my day job is helping companies position themselves in the cloud-native ecosystem so that their product’s value is obvious to end-users. I started this podcast because organizations embark on the cloud naive journey for business reasons, but in general, the industry doesn’t talk about them. Instead, we talk a lot about technical reasons. I’m hoping that with this podcast, we focus more on the business goals and business motivations that lead organizations to adopt cloud-native and Kubernetes. I hope you’ll join me.

Emily: Welcome to The Business of Cloud Native, my name is Emily Omier. I'm your host, and today I'm chatting with Ben Golub. Ben, thank you so much for joining us.

Ben: Oh, Thank you for having me.

Emily: And I always like to just start off with having you introduce yourself. So, not only where you work and what your job title is, but what you actually spend your day doing.

Ben: [laughs]. Okay. I'm Ben Golub. I'm currently the executive chair and CEO of Storj Labs, which is a decentralized storage service. We kind of like to think of it as the Airbnb of disk drives, But probably most of the people on your podcast who, if they're familiar with the, sort of, cloud-native space would have known me as the former CEO of Docker from when it was released up until a few years ago. But yeah, I tend to spend my days doing a lot of stuff, in addition to family and dealing with COVID, running startups. This is now my seventh startup, fourth is a CEO.

Emily: Tell me a little bit, like, you know, when you stumble into your home office—just kidding—nobody is going to the office, I know. But when you start your day, what sort of tasks are on your todo list? So, what do you actually spend your time doing?

Ben: Sure. We've got a great team of people who are running a decentralized storage company. But of course, we are decentralized in more ways than one. We are 45 people spread across 15 different countries, trying to build a network that provides enterprise-grade storage on disk drives that we don't own, that are spread across 85 different countries. So, there's a lot of coordination, a lot of making sure that everybody has the context to do the right thing, and that we stay focused on doing the right thing for our users, doing the right thing for our suppliers, doing the right thing for each other, as well.

Emily: One of the reasons I thought it’d be really interesting to talk with you is that I know your goal is to, sort of, revolutionize some of the business models related to managing storage. Can you talk about that a little bit more?

Ben: Sure. Sure. I mean, obviously, there's been a big trend over the past several years towards the Cloud in general, and a big part of the [laughs] Cloud is storage. Actually, AWS started with S3, and it's a $90 billion market that's growing. The world's going to create enough data this year to fill a stack of CD-ROMs, to the orbit of Mars and back. And yet prices haven't come down, really, in about five years, and the whole market is controlled by essentially three players, Microsoft, Google, in the largest, Amazon, who also happen to be three of the five largest companies on the planet.

And we think that data is so critical to everything that we do that we want to make sure that it doesn't stay centralized in the hands of a few, but that we, sort of, create a more, sort of, democratic—if you will—way of handling data that also addresses some of the serious privacy, data mining, and security concerns that happen when all the data is held by only a few people.

Emily: With this, I'm sure you've heard about digital vegans. So, people who try to avoid all of the big tech giants—

Ben: Right, right.

Emily: Does this make it possible to do that?

Ben: Well, so we're more of a back end. So, we're a service that people who produce-consumer-facing services use. But absolutely, if somebody—and we actually have people who want to create a more secure way of providing data backup, more secure way of enabling data communications, video sharing, all these sorts of things, and they can use us and service those [laughs] digital vegans, if you will.

Emily: So, if I'm creating a SaaS product for digital vegans, I would go with you?

Ben: I would hope you’d consider us, yeah. And by the way, I mean, also people who have mainstream applications use us as well. I mean, so we have people who are working with us who may have sensitive medical data on people, or people who are doing advanced research into areas like COVID, and they're using us partially because we're more secure and more private, but also because we are less likely to be hacked. And also because frankly faster, cheaper, more resilient.

Emily: I was just going to ask, what are the advantages of distributed storage?

Ben: Yeah. We benefit from all the same things that the move towards cloud-native in general benefits from, right? When you take workloads, and you take data, and you spread them across large numbers of devices that are operated independently, you get more resilience, you get more security, you can get better performance because things are closer to the edge. And all of these are benefits that are, sort of, inherent to doing things in a decentralized way as opposed to a centralized way. And then, quite frankly we’re cheaper. I mean, because of the economics and doing this this way, we can price anywhere from a half to a third of what the large cloud providers offer, and do so profitably for ourselves.

Emily: You also offer some new revenue models for open-source projects. Can you talk about that a little bit more?

Ben: Sure, I mean, obviously I come from an open-source background, and one of the big stories of open-source for the past several years is the challenges for open-source companies in monetizing, and in particular, in a cloud world, a large number of open-source companies are now facing the situation where their products, completely legally but nonetheless, not in a fiscally sustainable way, are run by the large cloud companies and essentially given away as a loss leader. So, that a large cloud company might take a great product from Mongo or Redis, or Elastic, and run it essentially for free, give it away for free, not pay Mongo, Elastic, or Redis. And the cloud companies monetize that by charging customers for compute, and storage, and bandwidth. But unfortunately, the people who've done all the work to build this great product don't have the opportunity to share in the monetization.

And it makes it really very hard to adopt a SaaS model for the cloud companies, which for many of them is really the best way that they would normally have for monetizing their efforts. So, what we have done is we've launched a program that basically turns it on its head and says, “Hey, if you are an open-source project, and you integrate with us in a way that your users send data to us, we'll share the revenue back with you. And as more of your users share more data with us, we'll send more money back to you.” And we think that that's the way it should be. If people are building great open-source projects that generate usage and revolutionize computing, they should be rewarded as well.

Emily: How important is this to the open-source community? How challenging is it to find a way to support an open-source project?

Ben: It's critical. I mean, if you look at the most—I’d start by saying two-thirds of all cloud workloads are open-source, and yet in the $180 billion cloud market, less than $5 billion [unintelligible] going back to the open-source projects that have built these things. And it's not easy to build an open-source project, and it takes resources. And even if you have a large community, you have developers who have families, or [laughs] need to eat, right? And so, as an open-source company, what you really want to be able to do is become self-sustaining. And while having contributions is great, ultimately, if open-source projects don't become self-sustaining, they die.

Emily: A question just, sort of, about the open-source ethos: I mean, how does the community about open-source feel about this? It is obvious developers have to eat just like everybody else, and it seems like it should be obvious that they should also be rewarded when they have a project that's successful. But sometimes you hear that not everybody is comfortable with open-source being monetized in any way. It's like a dirty word.

Ben: Yeah. I mean, I think [unintelligible] some people who object to open-source being monetized, and that tends to be a fringe, but I think there's a larger percentage that don't like the notion that you have to come up with a more restrictive license in order to monetize. And I think unfortunately a lot of open-source companies have felt the need to adopt more restrictive licenses in order to prevent their product being taken and used as a loss leader by the large cloud companies. And I guess our view is, “Hey, what the world doesn't need is a different kind of license. It needs a different kind of cloud.”

And that's, and that's what we've been doing. And I think our approach has, frankly, gotten a lot of enthusiasm and support because it feels fair. It's not, it's not trying to block people from doing what they want to do with open-source and saying, “This usage is good, this is bad.” It's just saying, “Hey, here's a new viable model for monetizing open-source that is fair to the open-source companies.”

Emily: So, does Storj just manage storage? Or, where's the compute coming from?

Ben: It's a good question. And so, generally speaking, the compute can either be done on-premise, it can be done at the end. And we’re, sort of, working with both kinds. We ourselves don't offer a compute service, but because the world is getting more decentralized, and because, frankly, the rise of cloud-native approaches, people are able to have the compute and the storage happening in different places.

Emily: How challenging is it to work with storage, and how similar of an experience is it to working with something like AWS for an end-user? I just want to get my app up.

Ben: Sure, sure. If you have an S3 compatible application, we're also S3 compatible. So, if you've written your application to run on AWS S3, or frankly, these days most people use the S3 API for Google and Microsoft as well, it's really not a big effort to transition. You change a few lines of code, and suddenly, the data is being stored in one place versus the other. We also have native libraries in a lot of different languages and bindings, so for people who want to take full advantage of everything that we have to offer, it's a little bit more work, but for the most part, our aim is to say, “You don't have to change the way that you do storage in order to get a much better way of doing storage.”

Emily: So, let me ask a couple questions just related to the topic of our podcast, the business of cloud-native. What do you think are the reasons that end users decide to go for cloud-native?

Ben: Oh, I think there are huge advantages across the board. There are certainly a lot of infrastructural advantages: the fact that you can scale much more quickly, the fact that you can operate much more efficiently, the fact that you are able to be far more resilient, these are all benefits that seemed to come with adopting more cloud-native approaches on the infrastructure side if you will. But for many users, the bigger advantages come from running your applications in a more cloud-native way. Rather than having a big monolithic application that's tied tightly to a big monolithic piece of hardware, and both are hard to change, and both are at risk, if you write applications composed of smaller pieces that can be modified quickly and independently by small teams and scale independently, that's just a much more scalable, faster way to build, frankly, better applications. You couldn't have a Zoom, or a Facebook, or Google search, or any of these massive-scale, rapidly changing applications being written in the traditional way.

Emily: Those sound kind of like engineering reasons for cloud-native. What about business reasons?

Ben: Right. So, the business reasons [unintelligible], sort of, come alongside. I mean, so when you're able to write applications faster, modify them faster, adapt to a changing environment faster, do it with fewer people, all of those end up having real big business benefits. Being able to scale flexibly, these give huge economic benefits, but I think the economic benefits on the infrastructure side are probably outweighed by the business flexibility: the fact that you can build things quickly and modify them quickly, and react quickly to changing environment, that’s [unintelligible]. Obviously, again, you use Zoom as an example. There's this two-week period, back in March, where suddenly almost every classroom and every business started using Zoom, and Zoom was able to scale rapidly, adapt rapidly, and suddenly support that. And that's because it was done in a more—in a cloud-native way.

Emily: I mean, it's interesting, one of the tensions that I've seen in this space is that some people like to talk a lot about cost benefits. So, we're going to move to cloud-native because it's cheap, we're going to reduce costs. And then there's other people that say, well, this isn't really a cost story. It's a flexibility and agility, a speed story.

Ben: Yeah, yeah. And I think the answer is it can be both. What I always say, though, is cloud-native is not really a destination, it's a journey. And how far we go along with that path, and whether you emphasize the operational side versus—or the infrastructural side versus the development side, it sort of depends on who you are, and what your application is, and how much it needs to scale.

And it's absolutely the case that for many companies and applications if they try to look like Google from day one, they're going to fail. And they don't need to because it’s—the way you build an application that's going to be servicing hundreds of million people is different than the way you build an application, there's going to be servicing 50,000 people.

Emily: What do you see is that some of the biggest misconceptions or mistakes that people make on this journey?

Ben: So, I think one is clearly that they knew it as an all or nothing proposition, and they don't think about why they're going on the journey. I think a second mistake that they often make is that they underestimate the organizational change that it takes to build things in the cloud-native way. And obviously, the people, and how they work together, and how you organize, is as big transition for many people as the tech stack that you’d use. And I think the third is that they don't take full advantage of what it takes to move a traditional application to run it in a cloud-native infrastructure. And you can get a lot of benefits, frankly, just by containerizing or Docker-izing a traditional app and moving it online.

Emily: What about specifically related to storage and data management? What do you think are some misconceptions or pitfalls?

Ben: Right. So, I think that the challenge that many people have when they deal with storage is that they don't think about the data at rest. They don't think about the security issues that are inherent in having data that can be attacked in a single place, or needs to be retrieved from a single place. And part of why we built Storj, frankly, is a belief that if you take data and you encrypt it, and you break it up into pieces, and you distribute those pieces, you actually are doing things in a much better way that's inherent, that you're not dependent on any one data center being up, or any one administrator doing their job correctly, or any password being strong.

By reducing the susceptibility to single points of failure, you can create an environment that's more secure, much faster, much more reliable. And that's math. And it gets kind of shocking to see that people who make the journey to cloud-native, while they're changing lots of other aspects of their infrastructure and their applications, repeating the same mistakes that people have been making for 30 years in terms of data access, security, and distribution.

Emily: Do you think that that is partially a skills gap?

Ben: It may be a skills gap, but it’s also, frankly, there's been a dearth of viable other options. And I think that—we frequently when I'm talking with customers, they all say, “Hey, we've been thinking about being decentralized for a while, but it just has been too difficult to do.” Or there have been decentralized options, but they're, sort of, toys. And so, what we've aimed to do is create a decentralized storage solution that is enterprise-grade, is S3 compatible, so it's easy to adopt, but that brings all the benefits of decentralization.

Emily: I'm also just curious because of the sort of organizational changes that need to happen. I mean, everybody, particularly in a large organization, is going to have these super-specific areas of expertise, and to a certain extent, you have to bring them all together.

Ben: You do. Right. You do have to. And so I'm a big believer in you pick pilot projects that you do with a small team, and you get some wins, and nothing helps evangelize change better than wins. And it's hard to get people to change if they don't see success, and a better world at the end of the tunnel.

And so, what we've tried to do, and what I think people doing in the cloud-native journey often do, is you say, “Let's take a small low-risk application or small, low-risk dataset, handle it in a different way, and show the world that it can be done better,” right? Or, “Show our organization that it can be better.” And then build up not only muscle memory around how you do this, but you build up natural advocates in the organization.

Emily: Going back to this idea of costs, you mentioned that Storj can reduce costs substantially. Do you think a lot of organizations are surprised at how much cloud storage costs?

Ben: Yes. And unfortunately, it's a surprise that comes over time. I mean, you… I think the typical story if you get started with Cloud. And there's not a lot of large upfront costs when your usage is low. So, yeah, so you start with somebody pulling out their credit card and building their pilot project, and just charging themselves directly to charging themselves directly to—you know, charging their Amazon, or their Google, or their Microsoft directly to their credit card, then they move to paying through a centralized organization.

But then as they grow, suddenly, this thing that seemed really low price becomes very, very expensive, and they feel trapped. And data, in particular, has this—in some ways, it grows a lot faster than compute. Because, generally speaking, you're keeping around the data that you've created. So, you have this base of data that grows so slowly that you’re creating more data every day, but you're also storing all the data that you’ve had in the past. So, it grows a lot more exponentially than compute, often. And because data at rest is somewhat expensive to move around, people often find themselves regretting their decisions a few months into the project, if they're stuck with one centralized provider. And the providers make it very difficult and expensive to move data out.

Emily: What advice would you have to somebody who's at that stage, at the just getting started, whipping out my credit card stage? What do you do to avoid that sinking feeling in your stomach five months from now?

Ben: Right. I mean, I guess what I would say is that don't make yourself dependent on any one provider or any one person. And that's because things have gotten so much more compatible, and that's on the storage side by the things that we do, on the compute side by the use of containers and Docker. You don't need to lock yourself in, as long as you're thoughtful at the outset.

Emily: And who's the right person to be thinking about these things?

Ben: That's a good question. So, you know, I’d like to say the individual developer, except developers for the most part, they have something that they want to build, [laughs] they want to get it built as fast as possible and they don't want to worry about infrastructure. But I really think it's probably that set of people that we call DevOps people that really should be thinking about this, to be thinking not only how can we enable people to build and deploy and secure faster, but how can we build and secure and deploy in a way that doesn't make us dependent on centralized services?

Emily: Do you have other pieces of advice for somebody setting out on the “Cloud journey,” in quotes, too basically avoid the feeling, midway through, that they messed up.

Ben: So, I think that part of it is being thoughtful about how you set off on this cloud journey. I mean, know where you want to end up, I think this [unintelligible]. You want to set off on a journey across the country, it's good to know that you want to end up in Oregon versus you want to end up in Utah, or Arizona. [unintelligible] from east to west, and making sure your whole organization has a view of where you want to get. And then along the way, you can say, “You know what? Let’s course-correct.”

But if you are going down on the cloud journey because you want to save money, you want to have flexibility, you don't want to be locked in, you want to be able to move stuff to the edge, then thinking really seriously about whether your approach towards the Cloud is helping you achieve those ends. And, again, my view is that if you are going off on a journey to the Cloud, and you are locking yourself into a large provider that is highly centralized, you're probably not going to achieve those aims in the long run.

Emily: And then again, who is the persona who needs to be thinking this? And ultimately, whose responsibility is it to make a cloud journey successful?

Ben: So, I think that generally speaking, a cloud journey past these initial pilots where I think pilots are often, it's a small team that are proving that things can be done in a cloud-native way, they should do whatever it takes to prove that something can be done, and get some successes. But then I think that the head of engineering, the Vice President of Operations, the person who's heading up DevOps should be thoughtful, and should be thinking about where the organization is going, from that initial pilot into developing the long-term strategy.

Emily: Anything else that you'd like to add?

Ben: Well, these are a lot of really good questions, so I appreciate all your questions and the topic in general. I guess I would just add, maybe my own personal bias, that data is important. The cloud is important, but data is really important. And as, you know, look at the world creating enough data this year to fill a stack of CD-ROMs, to the orbit of Mars and back, some of that is cat videos, but also buried in there is probably the cure to COVID, and the cure for cancer, and a new form of energy. And so, making it possible for people to create, and store, and retrieve, and use data in a way that's cost-effective, where they don't have to throw out data, that is secure and private, that's a really noble goal. And that's a really important thing, I think, for all of us to embrace.

Emily: Just a couple of final questions. The first one, I just like to ask everybody, what is your favorite can't-live-without software engineering tool?

Ben: Honestly, I think that collaboration tools, writ large, are important. And whether that's things like GitHub, or things like video conferencing, or things like shared meeting spaces, it's really the tools enable groups of people to work together that I think are the most important.

Emily: Where can people connect with you or follow you?

Ben: Oh, so I'm on Twitter, @golubbe, G-O-L-U-B-B-E. And that's probably the best place to initially reach out to me, but then I [blog], and I'm on GitHub as well. I'm not that great [unintelligible].

Emily: Well, thank you so much for joining us. This was a great conversation.

Ben: Oh, thank you, Emily. I had a great conversation as well.

Emily: Thanks for listening. I hope you’ve learned just a little bit more about The Business of Cloud Native. If you’d like to connect with me or learn more about my positioning services, look me up on LinkedIn: I’m Emily Omier—that’s O-M-I-E-R—or visit my website which is emilyomier.com. Thank you, and until next time.

Announcer: This has been a HumblePod production. Stay humble.

View Details

The conversation covers:

  • Josh’s role as CTO of Fugue, a leading cloud security and compliance provider for engineers.
  • The difference between cloud security and data center security — and why old school approaches to security don’t work in the cloud.
  • How engineers and security specialists can best communicate with business leaders about how to approach security, and how Fugue can help.
  • Who should be the person in charge of setting up Fugue, running reports, and communicating results across an oragnization.
  • The people who tend to lose their job when a cloud security breach occurs.
  • Why cloud security requires organizational change, and how companies are adapting to prevent issues.
  • The importance of upskilling employees and making sure they have the appropriate knowledge to solve cloud challenges.
  • Why the cloud has the possibility to be more secure than a data center. Josh also talks about cloud perception, and why some are still viewing the cloud as scarier than the data center.
  • What Joshn considers to be the most effective hacking strategies for cybercriminals.
  • The relationship between security and compliance, and how organizations should approach that relationship.
  • Why there is no such thing as a perfect security posture.

Links

  • Fugue: https://www.fugue.co/
  • Customer write-up on G2: https://www.g2.com/products/fugue/reviews/fugue-review-4269523
  • Twitter: https://twitter.com/joshstella
  • LinkedIn: https://www.linkedin.com/in/josh-stella-949a9711/
  • Fugue Blog: https://www.fugue.co/blog
  • Fugue Masterclass: https://resources.fugue.co/cloud-security-masterclass-registration
  • Fugue Office Hours: https://resources.fugue.co/cloud-infrastructure-security-office-hours

Transcript

Emily: Hi everyone. I’m Emily Omier, your host, and my day job is helping companies position themselves in the cloud-native ecosystem so that their product’s value is obvious to end-users. I started this podcast because organizations embark on the cloud naive journey for business reasons, but in general, the industry doesn’t talk about them. Instead, we talk a lot about technical reasons. I’m hoping that with this podcast, we focus more on the business goals and business motivations that lead organizations to adopt cloud-native and Kubernetes. I hope you’ll join me.

Emily: Welcome to The Business of Cloud Native. I'm Emily Omier, your host, and today I'm chatting with Josh Stella. Josh, thanks so much for joining us.

Josh: Well, Emily, thanks so much for having me.

Emily: Of course. I always like to start the same. Can you just introduce yourself and your company, and tell me a little bit about what the company does, and then also what you do?

Josh: Sure. So, Fugue does cloud security for public cloud providers like AWS, and Azure, and Google. Prior to founding Fugue, I worked at AWS as a principal solutions architect primarily focused on national security; Department of Defense, and similar things. My background is I'm a programmer and I'm a software architect, and I've kind of lived between national security kinds of work and high tech in startups. And so what Fugue does is we’ll tell you all about the security posture of your cloud environments, and teach you where you have weaknesses that hackers can exploit; we help you close those, and then we can actually keep things from having those misconfigurations going forward. So, that's a little bit about us. If you're a developer, you can use our forever free developer version, and we work with a lot of enterprises folks like SAP, and big organizations, too.

Emily: So, were you involved with setting up the super-secret CIA cloud that AWS was involved in?

Josh: I was not personally. A very close colleague of mine was actually working very closely on that, but no, I was not directly involved in that.

Emily: Okay, you probably couldn't talk about it, even if you were so. [laughs].

Josh: No comment.

Emily: Anyway, I always like to ask also, what do you actually do? Like, you get up in the morning, presumably, you don't go to an office anymore, but—

Josh: Oh, true. True, yeah. Whether going to an office or not, my days are… so I started out founding the company with my co-founder, Andrew Wright. And for a while, I was the CEO when we were in the kind of R&D phase, but then I always intended to hire a really great CEO, which we did a couple of years ago, Phillip Merrick, and I became the CTO. And there are different kinds of CTO.

My main functions are, like, I get up in the morning, I go read the news about any breaches in Cloud that have happened, and then I try to recreate them whenever possible, if there's enough information, because the attack vectors on Cloud are completely different than in the data center, and are inobvious to folks. So, when you read about a breach, and you see that they use the identity and access management service almost like a network, to get to S3, that's really interesting and it's really important so that Fugue can protect our customers. So, I spent a fair amount of time doing that. I do work every day with the product team. Occasionally, I will weigh in fairly strongly on an engineering topic, but a lot of times our engineers are just very, very good and we've hired experts and all their areas so I work with them, but it's usually just to give advice and some guidance.

And I do a fair amount of writing, and I do a fair amount of teaching classes online: we have a masterclass series on Cloud security that has been very well received. And then the research I do into how cloud exploits are actually being done by recreating those in my own environments, I use those both in the classes and of course, Fugue as our product can then have protections built-in against them. So, I’d say that's a lot of what I do.

Emily: I wanted to ask a little bit more about this difference between cloud security and data center security. Can you go into that a little bit more? And then also, what do people miss in that difference?

Josh: Okay, so I'm going to start at the prosaic and kind of go to the sublime a little bit, but the most simple way to think about the difference is in the data center days, you really had a network perimeter. So, you've got a big pile of servers, they're racked and there are switches that that connect them together, and then there's this layer of security at the, kind of, perimeters of the network where the data center network connects to, whether it's the corporate network, or another data center, or the internet. And that kind of perimeter defense slash defense in-depth idea meant when you were talking about data center security, the primary things you were thinking about were, “What's happening on my network?” And, “Are the servers—or the physical devices that are actually running compute stuff—are they secure?” Well, it turns out in the Cloud, almost none of that matters that much, and the reason is—so I think Gartner recently said, like, over 90 percent of cloud exploits are due to misconfigured cloud services.

So, that the data center, you had, like, piles of baryonic matter. You had actual servers in racks, you would replace them every three years, maybe five years on a recap cycle. And so it was moving slowly, and to get to those things, you had to kind of penetrate those layers of network perimeter and defense in depth. Well, on the Cloud if you stand up an S3 bucket, for example—and the press loves to pick on S3 breaches because well, a lot of people store data in S3, but very often these breaches are much more complex than just a misconfigured S3 bucket—but Gartner said the vast majority—and this is what we see, and what I saw at AWS—of hacker success in Cloud is looking for misconfigured cloud services, and then exploiting those. So, wherein the old days you might have been really concerned: “Do I have a bunch of packets coming in from an IP address that's known to have botnet activity on it?”

And if so, I'll try to shut off that flow of packets, and, “Are my server operating systems patched?” And things like that. You definitely still want to keep your patch levels up, but a more typical cloud exploit would be something like finding API keys in an insecure place so that the attacker can modify your cloud infrastructure, and goes in and steals the backup of a database, or installs crypto mining software inside your infrastructure. So, it's a really big topic, but you really have to think about it very differently. You have to think about it as a software engineer more than as a security engineer. You have to think about it less as, “I'm protecting things,” and more as, “I'm configuring them properly. I'm making the Cloud safe by configuring it properly.” This is why we teach masterclasses. There's a very long list of ways to get those things wrong.

Emily: Do you think in general, people are aware of these differences, and this idea that you have to think about security differently? Like, how much is that percolating through to people who are not in—they're not living and breathing this cloud computing space?

Josh: It has not dawned on as many people as I wish it had. [laughs]. There are still a lot of folks who think using old school approaches to security will work on Cloud. And there are a number of problems with that. One is that the skills needed to do this stuff properly on Cloud are more like software engineering, engineering skills, and less analyst skills.

So, when you can configure all of your compute resources via APIs, which is how the Cloud works, you're going to automate—I can build a global network while I'm talking to you by typing in the background. You could not do that in the data center. So, I don't see enough people being aware, not fully. And the worst thing that I see—well, and it's particularly bad if you don't understand what the threat is and what the attack surfaces are because, guess what, the attackers are all automated and very clever, and they will get you—but the next bad thing I see is when people try to force the Cloud to act like a data center in order to use the old ideas about how to do security. And that removes all the advantages of the Cloud, and it also never works.

Emily: That's interesting. Obviously, this podcast is about the business of cloud and cloud-native, and security is sort of ultimately a business problem.

Josh: Yep.

Emily: How do people talk about security when they're talking, not just inside the engineering department, but also with business leaders, and how can engineers—or security specialists—how can they communicate, “This is the way we need to approach security. This is why. This is what we need to do and why?” How do people make sure there's not some stuff that gets lost in translation?

Josh: What I recommend is that—and this is why we—it's one of the reasons why we built Fugue, is we give people tools for doing this—what I would recommend is kind of base camp one on climbing the Cloud security mountain is just understanding your current security posture; understanding whether what you have built in the Cloud is safe. And I can pretty much guarantee to every listener that, you know, we had a new customer write a nice write-up on Fugue in G2 saying, “Fugue is going to hurt your feelings the first time you run it.” We're going to tell you about a whole bunch of stuff that should scare you, and we're going to present that visually, and that's something you can take to a boss. You can say, “Hey, I used this free tool and it says we have, like, 90 things configured in ways that hackers exploit. We should go fix those, and by the way, we should keep keeping track of this.” And so I think presenting the information, rather than it being vague, quantifying it and having evidence, in my experience—in general in life around business decisions—is much better than having an opinion or winning an argument. So, get some data, show it to your boss, and go start fixing stuff. [laughs].

Emily: So, my next question is who should do this? I mean, I think there's often a question related to security, and particularly you were just saying part of the issue with cloud security is that there's a little bit of a shift of responsibility, but who should be the one that's setting up Fugue, and running these reports, and talking to their boss about this problem?

Josh: Yeah. So, it varies because organizations are struggling with where these things should live. What I believe is the concern with security, as you pointed out, is fundamentally a business concern. It's: “are my systems doing what is intended for whom it is intended, and only that?” So, that has to begin with the folks that are building the systems.

So, very often at Fugue, what we'll see with the more sophisticated customers we have, the DevOps team will want this stuff baked in really early in the software development lifecycle, and they might even be doing that kind of independently of the security team, but the security team usually also has a role as does the compliance team. So, what we believe is that this should be implemented throughout the entire software development lifecycle from when people are writing their infrastructure’s code or building infrastructure in the development environments, all the way through to monitoring and production, and doing audit reports for SOC 2 compliance or whatever. Generally, with Fugue—depending on the organization—we will either find folks who really care about what we do in the DevOps team, or in the security team. And the way we like to talk about it is any engineer who is concerned with security—whatever their title and role—can benefit from understanding cloud security and being more effective at it.

Emily: When there is a cloud security breach, who loses their job?

Josh: Uh, well, quite a few CSOs. If it's the kind that hits the Wall Street Journal, usually the CSO is not going to be around. And if you can agree with what I'm saying, which is that building stuff correctly and not misconfiguring cloud is the most critical thing to get right for cloud security CSOs don’t generally have authority over how people build software. And I think this is a disconnect that really needs to be addressed in every organization. If you get breached, you might say, “Well, that's the chief information security officer’s job to keep those from getting breached.”

In the old days, that meant the security organization adding in those layers of perimeter defenses and trying to capture things. Well, now it means, you know, Global 2000 or Fortune 500, there are thousands of developers, any one of whom might be building a new network right now, or a whole new application and all of its attendant infrastructure. So, I think we're going through a period where that shift in technology has created a shift in responsibilities, or maybe a shift in the ability to do the right things and where that needs to happen, but the old ways of thinking about how information technology was built and secured aren’t helping. So, the CSO gets fired, but it's usually not their fault is my opinion. [laughs].

Emily: What would you think—I mean, everything I talked about with moving to Cloud and moving to cloud-native, there's all these technical changes, but there's also all these organizational changes. How does cloud security require organizational change? And how successful do you see companies being at adapting not just tech, but also organizations?

Josh: Well, I mean I think that is the important part. The tech is there: we built Fugue, there's other tools out there, there's other things you can use, there is great technology for this. The struggle is with the organization and how to implement it, and how to operationalize it, and build it into workflows. So, I think that as far as how well people are doing with it: it's highly variable, and it's even highly variable within organizations. So, you might find a sort of pocket of people who are very clued in to how you should be thinking about this and dealing with it in the same company where another group is still thinking, our firewalls and our intrusion detection systems are good enough, or are more important than they are.

So, it really comes down to whether you are thinking in a cloud-native way. And I think that the Cloud is not a pile of remote data centers. It is a global distributed computer that you can program and configure. And that's really what we're talking about when you build an air-quotes, “network,” Amazon aren't running around and plugging wires into switches. That's just a configuration. It's just a configuration of that big distributed computer, and so we have to think about this from a software engineering perspective.

Now, the good news is there, that—well as it relates to the organization, a lot of security organizations don't have a lot of developers in them, and so this looks confusing and scary. And that always creates challenges. If people are intimidated by something or it's out of their comfort zone, that actually creates organizational friction. But I'm here to tell you, the Cloud is potentially—if done well—Cloud is the most secure way to do computing ever invented by humans, and for a really simple reason: you can control it all through APIs. Now, that means I can build a global network in five minutes and you might not notice, but it also means I can write programs that are constantly aware of everything and know how to get things right. So, it is a big organizational challenge, a lot of folks that are trying to, kind of, adapt their traditional data center teams may not be aware that you probably have people on your teams that are really trying their level best but don't have the most appropriate skills for the problem that is now required to be solved.

Emily: Skilling up is always a big challenge, as is reorganizing and reconceptualizing how you work and how you work together.

Josh: Oh, yeah. Yeah, absolutely. And a lot of the—it gets really complex because organizational structure has a whole lot more inertia than technology, and for lots of reasons: somebody has been here for 15 years and wants to get promoted, or hasn't been promoted and has a whole team and a budget, and who's going to pay for this, and who's going to chip into this, and which managers need to change their ideas about what they own and are responsible for? And I'm not trying to be pejorative here. That's all real stuff.

The Cloud has effectively turned security on it’s head. It used to be the infrastructure and security teams would go build secure environments, and then application developers would deploy into those environments. And now all of a sudden, application developers run a script or a program, and it builds the “infrastructure”—air quotes, again. We're really configuring a big, existent thing—but all of a sudden, in five minutes, you have what it would have been weeks and months of procurement, and discussion, and controls on what was being built. So, when you have a change that radical, you probably can expect—you should expect—that a lot of people are going to be challenged by that and probably won't tell you that they are because people don't like to admit when they don't understand something.

So, skilling people up is really, really important. What we try to do in the masterclass series is give people new ways to think about these topics. I can't teach you everything I know about cloud security in a series of classes that are only an hour every other week or something, which is about what we do, but I can tell you how to take apart the problems and really think about them. And guess what? Even if you don't already know how to do that, you can learn it, and it's interesting, and it's fun.

Emily: I wanted to go back to something else that you said about Cloud having the possibility to be more secure than a data center. How do you think that figures—if at all—into organizations’ decision to move into the Cloud? I mean, how aware are they that this actually could be more secure?

Josh: Highly variable. Some understand that. Most don’t. Most are still viewing Cloud as, kind of, scarier than the data center because their mindset hasn't shifted. So, for example, with Fugue, you can tell Fugue, “Tell me everywhere I have a dangerous misconfiguration.” And we will show that to you on an automatically generated map, like a Google map of your infrastructure. We'll show you physically, visually where that is.

Well, how would you do that, even that simple thing from a Cloud—Fugue perspective, in the data center? You would have to typically do what's called a data call, and get a bunch of human beings to answer questions and plug data into spreadsheets, and when I worked in national security, we would get these data calls, and you'd be given, like, a week. Well, Fugue can do that every five minutes, in about five minutes. So, even out of the gate, the fact that you can use computer software to interrogate the other computer software means you can automate it, and you can be much more thorough. Humans, also, are terrible at keeping lists of details. Computers are great at it, so we can use them for that.

But then as you extend further—and we have this diagram of climbing a mountain of cloud security—so the very first base camp, sort of, at the foot of the mountain is just understanding what you have and what's wrong. But as you go up, you want to start doing some other things: you want the system to automatically tell you any time something changes, and any time something changes in a dangerous way. And then with Fugue, we've, kind of, taken this to its logical conclusion—and we're unique in this regard. This is actually what we started with when we first founded the company, and why we did so much R&D—you can tell Fugue, if anything is misconfigured, automatically heal it back to a known good configuration. Fully self-healing infrastructure is something that was totally elusive for decades in the data center because there weren't consistent APIs over these things. So, not a lot of people are ready to think that way, but that's where this is all going.

Emily: So, is this kind of like how people are more afraid to fly in an airplane because they're not in control, but they have no problem getting into their car, even though you're, like, dramatically more likely to die in a car accident.

Josh: That is an awesome analogy and I'm going to steal it. Yeah, it really is a lot like that. I think that in the history of automated systems in computer security, a lot of claims were made that couldn't be backed up with tech because we did not have these very consistent APIs that the Clouds have. If you were trying to do automated security in a data center, you probably got burned. So, in a way—and I don't want to strain your analogy too much, but I might a little bit—early flight was not so safe. Until you really get into the ’60s and ’70s, you should probably have been nervous getting on an airplane in 1937 or 1945, but by the time we really had figured it out, it became tremendously safer—as you point out—than driving a car.

And I think that folks are just starting to realize it. But you have to trust the system, you have to trust the automation, and that takes time and building trust. But I'm here to tell you, this can now be done in a way that will work highly effectively and it is the future. The bad guys are all automated, okay? The hackers are running scripts, and programs and botnets constantly to find anything you have facing the internet that has any exploit they know how to use. So, they just get up in the morning and have a cup of coffee and look through a log of all of your and everyone else's stuff that you got something wrong on, and they point another program that and exploit it. Well, if you're not using automation and they are, you're doomed.

Emily: Yeah, I wanted to ask, so if you went over to The Dark Side and decided to be a hacker, what would you do? What do you think are the most effective hacks? You don't have to go in detail, of course, but I'm just curious what you've learned.

Josh: Sure. I mean, there's a lot of ways to do this stuff. I mentioned two of the most common ways are finding—so if everything is driven by APIs, if you can get a hold of the API keys—essentially the login information, the ability to access APIs—if you can get a hold of those, then whoever's stuff is using those APIs you can now get to. And so you see a lot of that, folks using bad engineering practices and putting keys in source code, and it's showing up in GitHub, and people searching GitHub for keys. Or doing things like having unencrypted backup snapshots of file systems sitting out there and in those file system backups there being keys. That's something bad guys do a lot.

You should be using things like IAM, and key rotation, and putting everything in KMS, and doing best practices there, and never ever, ever having API keys stored in source code or on a disk or anywhere that the bad guy might find, it even later. I mean, there was a breach last year of actually a cloud security company—not us—where that's exactly what the bad guys did. They found—I won't name names—but they found some keys, and then they managed to get into the company's cloud environment—and this is really a cool hack. Instead of—the keys they found gave them access to the production database, but they didn't go into it. Why didn't they go into it? Because they're probably monitoring the production database. But those same keys allowed the bad guys to stand up another database cluster with a backup of the production database. Like that is such a cloud-native hack, you would never have hackers—well, I can't say never, but it would be extraordinarily difficult to imagine hackers breaking into a data center and then standing up a new compute cluster to steal data from so that it wouldn't be monitored because that would be obvious.

But in Cloud, that's probably happening all day, every day where people are building stuff, and that's why something like Fugue becomes important. So, there's lots of ways you can go about it, and honestly, one of the most fun parts of my job is reading about those things and seeing how hackers are going about it because I'm telling you, they are more clever and more cloud-native than most of the people that are trying to prevent them coming in, and I'm often just very impressed and, kind of, I'm not happy anyone's data got stolen, but I can appreciate a good hack and there's some really clever ones out there. IAM relationships is another big one to look at. People still think of IAM—identity and access management stuff—as being identity. And it is, but now it's the identity of compute resources, and it forms a network that sidesteps the TCP/IP-based network, and so if you can surf that network, nobody's probably monitoring it and you can get a lot of places. So, I can't answer more succinctly than that, but those are a couple of ways.

Emily: Before we go, I wanted to ask you about the relationship between security and compliance. So, even if security is clearly a business problem, compliance is, like, even more into that business-y category. Can you just talk a little bit about the relationship between security and compliance and how organizations need to think about that relationship?

Josh: Well, it's a good question. So, historically—at least in the environments that I worked in—compliance and audit were, sort of, at the end of the development process as a sort of gating function. In the national security world, you need to get what's called an ATO, an authority to operate, and an ATO requires going through a NIST certification and accreditation against the 800-53 compliance family standard, and that's a big manual process with a human team, and Excel spreadsheets, and big documents and so on. And it makes sense. You do want to try to prevent bad stuff from happening.

Well, in Cloud—and then security was kind of different in that security were the folks who were doing things like monitoring firewall configurations and doing packet capturing at those perimeters that don't exist anymore, and doing intrusion detection, and making sure people's laptops didn't have unapproved applications on them. So, these worlds were a little separated. In Cloud, they're actually much more closely related, and they should be. And the beauty of that is, when you look at these compliance standards, like NIST 800-53, like SOC 2, like PCI, or GDPR, or HIPAA, or any of them—we cover a whole bunch of them—they're going to give you a whole lot of good advice from a security perspective. So, you can now, using automation and tools like Fugue, rather than having this dreaded audit come, or CNA process at the end of your development cycle that's going to add weeks to—of friction before you can deploy as you fix errors, you can just constantly be using those compliance standards to check your work in an automated way.

So, you build a little, you find out that, “Well, am I breaking any NIST rules?” You know, NIST is going to tell you—well that NIST implemented correctly, like in Fugue, Fugue’s going to tell you, “NIST says you have to have all your data be encrypted at rest everywhere,” And we're going to look across hundreds of cloud resource types and tell you if anywhere, you have data that's not being encrypted at rest. So, it really changes the game on Cloud because now compliance through automation, so it's not this manual audit at the end anymore, you can now have it completely automated and baked into the software development lifecycle instead of doing a week-long audit at the end. In five minutes, you get that feedback, and you can keep doing it iteratively. Compliance can actually become a massive help to getting things secure all the way through the lifecycle.

And in fact, I would point folks to a good friend of ours, was the guest star of a class about a week ago, and he's an expert on bringing cloud environments into compliance. And he taught a whole class on that, I'd recommend that. The final thing I'll say, though, is all those CIS benchmark, and NIST, and all those, there's lots of good stuff in there, but what we've learned from recreating these cloud hacks, is that the hackers are ahead of the compliance standards and the security teams. And so in Fugue, we bake in what we call Fugue Best Practices, and really what that is, is a collection of stuff that will tell you if you're vulnerable to the kind of hack that you read about in the news recently. And you're not necessarily going to get that—you won't get a complete picture of that with things like NIST, and CIS, and so on. However, they're awesome; they're going to tell you a lot. I hope that answered the question. I hope I answered the correct question there. [laughs].

Emily: Oh, absolutely. Well, first of all, it sounds like you can be completely compliant and still get hacked, but I think everybody knows that, you know, there's no such thing as a perfect security posture.

Josh: That is very true. The security posture is going to be helped by looking at compliance standards. There are other things we know are dangerous that are not in the compliance standards, and that's why we put those—and by the way, those things are actually in the totally free forever version of Fugue; you can go see if you're vulnerable to this because the compliance bodies are, kind of, slower-moving, and we can be faster. But then there's another third category, which is hackers doing stuff that no one predicted they would do.

And guess what? They're good at that. There's a reason why hacking used to mean a clever program, and now the means breaking into your stuff because they do it through cleverness, through deep technical expertise. So, you're not going to be able to predict what they're going to do, and that's their job. And therefore you have to employ other tactics than just compliance.

You have to employ things like drift detection: noticing if anything changed in the environment’s configuration. And again in Cloud, 90 percent of what you should care about is configuration of Cloud. Not logs, not packets going over networks. Those are leaky abstractions on Cloud. They don’t capture—there is no real perimeter, so you really have to be thinking about configuration.

And you want to use compliance standards, you want to use predictive rules, but then you also need to keep track of what's going on. So, for example, if I saw a compute instance change its IAM role association, I would immediately—if I got a notification that that happened, I would be immediately looking into that because that has the potential to be a devastating attack, and probably the biggest breach anyone has ever heard of, that's how it happened. So, we now predict that in Fugue Best Practices, but you really need to get things right from a security and compliance perspective, but then keep in mind that the hackers are going to do things that you're not predicting, and that's why we do drift detection and self-healing in Fugue because you just can't think of every bad thing they might do to you.

Emily: I think also part of what you're saying is just don't think of compliances as exclusively this hurdle that you have to jump through, but also think of it as almost like a tool that you can use, a set of best practices that you can use throughout the process.

Josh: Oh yeah, absolutely. I mean, that's the beautiful thing about what's happening now with Cloud. It’s stuff that used to be this onerous data call, and you have to fill out forms. You can just use—so, okay. I'm a programmer. I'm not a security engineer as a background, I'm a very security-focused programmer and software architect.

When I'm writing a program, I have tools that tell me where I'm being dumb. That’s, like, 80 percent of the job is your tools telling you where you've made errors, and then the other 20 percent is you catching the errors the tools weren't smart enough to find. So, for example, just to use a, kind of, goofy, trivial example, if I were to try to multiply your name times a date, if I have a decent programming compiler, or interpreter, or debugger, it's going to tell me, “You probably don't want to multiply a name by a date. You're probably not going to get a result that is sensible.” So, it's going to tell me where I've made a mistake.

With Fugue and similar technologies, that can now be done for security and compliance. And we use the compliance families to provide that guidance. So, in the same way that a programmer in the past would see, “I’ve made a cast error on types,” or something, now with using Fugue, Fugue will tell you, “Hey, you made a security error on that firewall rule.” With Cloud, and APIs, and automation tools like Fugue, security and compliance become highway builders, not tollbooth operators. They contribute to velocity rather than taking it away, and I think that's really exciting.

Emily: I just have a couple, sort of, last questions for you. The first one is what tool could you not live without, or I should say, do your job without?

Josh: Well, to be honest with you. The most important tool I use, other than things like web browsers that are just par for the course, is just having a great text editor. [laughs]. And programmers out there will understand why I'm saying that. And I got to say, I was an Emacs guy for, like, 20 some years, but VS Code is really, really good. I love what Microsoft's doing these days with the programming tools, and so I'll choose VS Code. I love it.

Emily: And then how can listeners connect with you, follow you, read more?

Josh: Oh, cool, yeah. So, on Twitter, I'm @joshstella. They can email me, I'm Josh, J-O-S-H@Fugue, F-U-G-U-E.co. I'm on LinkedIn. And if you ping me on any of those—I mean, the main thing I would suggest is keeping track of our blog, and the masterclass series, and we also do office hours. I mean, we take education really, really seriously at Fugue and trying to educate folks about what we have learned. And so hopefully people will find those things valuable. A lot of folks have.

Emily: Well, thank you so much, Josh, for joining us.

Josh: Well, yeah. Again, thanks for having me. I enjoyed the conversation, Emily.

Emily: Thanks for listening. I hope you’ve learned just a little bit more about The Business of Cloud Native. If you’d like to connect with me or learn more about my positioning services, look me up on LinkedIn: I’m Emily Omier—that’s O-M-I-E-R—or visit my website which is emilyomier.com. Thank you, and until next time.

Announcer: This has been a HumblePod production. Stay humble.

View Details

The conversation covers:

  • The difference between cloud computing and cloud-native, according to AJ
  • Whether it’s possible to have a cloud-native application that runs on-premise
  • The types of conversations that AJ has with customers, as VP of product. AJ also talks about the different types of customers that DigitalOcean serves.
  • How the needs of smaller teams tend to differ from the needs of enterprise users — and the challenges that smaller teams face when learning and implementing cloud-native applications.
  • Making decisions when using Kubernetes, and how it can be overwhelming due to the sheer number of choices that you can make.
  • Some of the main motivations that are driving smaller companies to Kubernetes. AJ also explains what he thinks is the best rationale for using Kubernetes.
  • Popular misconceptions about cloud-native and Kubernetes that AJ is seeing.
  • Why customers often struggle to make technology decisions to support their business goals.
  • AJ’s advice for businesses when making technology decisions.
  • Why startups are encouraged to start by using open source — and why open source wins in the end when compared to proprietary solutions.

Links

  • DigitalOcean: https://www.digitalocean.com/
  • Twitter: https://twitter.com/apurvajo
  • LinkedIn: https://www.linkedin.com/in/apurvajo/

Transcript

Emily: Hi everyone. I’m Emily Omier, your host, and my day job is helping companies position themselves in the cloud-native ecosystem so that their product’s value is obvious to end-users. I started this podcast because organizations embark on the cloud naive journey for business reasons, but in general, the industry doesn’t talk about them. Instead, we talk a lot about technical reasons. I’m hoping that with this podcast, we focus more on the business goals and business motivations that lead organizations to adopt cloud-native and Kubernetes. I hope you’ll join me.

Emily: Welcome to The Business of Cloud Native. I'm Emily Omier, your host, and today I'm chatting with AJ. AJ, can you go ahead and introduce yourself?

AJ: Hey, I'm AJ. I’m vice president of product for DigitalOcean. I've been with the company for about 15 months. Before that, I spent about a couple of decades with Microsoft. I was fortunate to work on Azure for the last decade, and I had the opportunity to build some cloud services with the company.

Emily: And thank you so much for joining us.

AJ: Thank you, thank you for having me.

Emily: I always like to start out by asking, what do you actually do? What does a day look like?

AJ: [laughs]. It’s an interesting question. So, yes, the day is usually all over the place depending on the priorities and things that are in motion for a given quarter or a week, per se. But usually, my days involve working with the team around the strategic initiatives that have been planned, driving clarity around different projects that I [unintelligible]. Mainly working with leadership on defining some of the roadmap for the product as well as the company. And yeah, and talking to lots of customers. That's something that I really, really enjoy. And every other day I have a meeting or two talking to our customers, learning from them, how they use our products and how can we get better.

Emily: I'm going to ask more about those conversations with customers because that's what I find really interesting. But first, actually, I wanted to start with another question. What do you see as the difference between cloud computing and cloud-native?

AJ: The difference essentially, in a way, the cloud computing is a much bigger umbrella around how we as a technology industry are enabling other businesses to bring their workload to a more scalable, more efficient, more secure environment versus trying to host, optimize, or do things by themselves. And the cloud-native, in a way, it's a subset of a cloud computing where not necessarily you always have to have existing workloads or something that is prior technology that has been already built and you're looking for a place to host. In a way, when you're building something out, new, greenfield apps and whatnot, you're starting from scratch, you're building your applications and solutions that are cloud-native by definition. They're built for Cloud; they're born in Cloud, and are optimizing the latest and the greatest innovations that are present and as future-looking to help you scale and succeed your business, in a way.

Emily: Do you think it's possible to have a cloud-native application that runs on-premise?

AJ: There's a lot of [laughs] innovations happening in pockets, and especially from the top providers to enable those scenarios. But at the end of the day, those investments are essentially driven to help people and companies, especially on the larger scale, to buy some time to completely move to the public cloud where the industry takes their time to come up with the compliance, security requirements and [unintelligible]. So, you'll start to see—you might have heard about some of the investments these top cloud providers are doing about allowing and bringing those similar stack and technologies that they are building in a public cloud to on-premise or running on their own data center, in a way. So, it is possible, in bits and pockets to start with a cloud-native to run, on-premise, but that customer segment and the target is very, very different than the ones that start in a public cloud first.

Emily: I want to switch to talking about some of the conversations that you have with customers. I really like to understand what end users are thinking. What would you say when you talk to customers? What's the thing that they're most excited about?

AJ: Right. So, it depends on what segment of customers you're speaking with, right? DigitalOcean serves a very different set of customers than a typical large cloud providers do. We're focused more on individual developers, small startups, or SMBs. Again, when I say SMBs, it's a broad term, when I say SMBs the S with [unintelligible].

So, we focus mainly on two to ten devs team, and smaller companies and whatnot. So, their requirements are very different; their needs are very unique compared to what I used to talk, back in my past life, with enterprise customers. Their requirements are very unique and different as well. So, what I hear from the customers that I speak with recently, and have been speaking with for last over a year, is how can I make my business that is [unintelligible] on a cloud? And what I mean by that is how do I build solutions that are simple, easy to understand, and where I'm focused on building software and not really worrying about the complexity of the infrastructure, at the same time, keep the price in control and very simple and predictable.

And that resonates really, really well. The tons and tons of customers that I spoke with recently, they moved from large cloud providers to our platform because their business was not viable on those cloud providers. And what I mean by that is you because of the complexity of how the [prime systems] product offering and the sticker shock they get, at the end of a month on a bill, it just does not make any business sense for them to keep on running on that cloud. So, they're looking for that kind of simplicity; they're looking for price predictability; they’re looking for something easy to get started, and cheap so they don't break the bank. I mean, those are some of the real common themes that I keep hearing from the business side of the customers, SMBs.

Then there's a set of different customers that I speak with. They are individual developers, they are students, they are people who are trying to learn technology. And what DigitalOcean has done great for our last eight, nine years is not only build this platform but build this great community of people who come to learn about technology. I would say more than 50 percent of our customers tell me how they love DO and they found DO because they were trying to learn so-and-so technology and they came and hit our tutorials. And they got to know about the company from the tutorials, they started learning, and it was very easy to transition to move into the platform because that's something they absolutely love.

Emily: When you think about the needs of these smaller teams, how would you contrast that with the enterprise users that you used to talk to more?

AJ: Yeah. Their needs are very different and very unique. They just need bare minimum basic things to get started with their application. They want to avoid all these complexity of hundred-plus services and trying to make the decision, which is the right one, which is the wrong one to go ahead with. So, they're not really looking for all kinds of bells and whistles and security requirement, or that feature that only one out of hundred customers will need or whatnot.

They're just looking for very minimum kind of a solution from the technology perspective. Something really simple, really easy to start with, not too many options out there that causes more confusion versus getting started quickly. So, simplicity is the biggest aspect that I see from this customer segment. And then second on the simplicity side is also simple pricing. For example, they would love to build applications and deploy and across the world, [across] regions, or [unintelligible] more regions to target different segments and customers. And they want to have the same pricing.

And bandwidth pricing is a great example. If you look at the other cloud providers, you end up paying different bandwidth pricing depending on which region you're deployed on. Unlike DO where the bandwidth pricing is flat across the world. So, that's really appealing to the customer set. Same thing goes with another product like storage, where you store your data, and when you are accessing the data, you need the predictability, you need to know how much it's going to cost you as your end-user usage pattern changes.

If you look at the top cloud provider, if you store your data in their storage, they will charge you for storage they will charge you hot data cold data, different API calls, you know, API calls for seconds and whatnot, and by the end of the day, your bill is very unpredictable compared to what you see at DO where you pay a flat monthly price, and none of those complexity comes in. So, that is a unique change that I see. And then, again, their requirements are very different, very, very simple. They don't need all the complexity. Something really simple to start with, including pricing.

Emily: This is interesting because most people, or… I should say, I. When I think of cloud-native, and I think even of cloud computing, simple is not the word that comes to mind.

AJ: Right.

Emily: How frustrating is this for these smaller teams to try to wrap their head around everything that they need to learn and understand in order to successfully use cloud-native applications?

AJ: It is very frustrating. And Kubernetes is a great example. It has a huge mindshare in the industry. It is a hugely popular technology that everybody's moving towards because it's the next big thing and the cool thing out there in the market, but the reality is something really, really complicated to learn, to understand. So, it is a super frustrating for people to not only learn the technology that is new and complicated but at the same time trying adapt to different versions of the similar implementation across different cloud because they were built to solve for a few specific customers in [unintelligible] that might not be suitable to them.

The one thing that I keep on hearing from the customers that are using our Kubernetes, they just love the fact how vanilla Kubernetes is offering it is. It's not too many bells and whistles. They have to really understand what this now means or what that now means, versus just, I understand the basic concept that I read from the tutorials, and I get what I'm seeing there. So, it is frustrating, at least on the cloud-native side, and the more layer of abstraction that you provide and make it simple, the onboarding becomes really, really exciting for them. But to my [unintelligible], to provide this simplicity to our end customers, we have to take on a crazy amount of complexity on our side, on the back end, and on the infrastructure. So, it's even more frustrating for our engineering to make sure they are delivering on the product requirements that my team comes up with because we want to keep things simple and straightforward for our customers.

Emily: A lot of people think one of the advantages of Kubernetes is being that it's infinitely extendable; it's infinitely flexible. And yet, obviously, when you have, sort of, infinite options, that means you have to make a bazillion different choices in order to just get something to work. How do your customers tend to approach those trade-offs? Do they want somebody to make the decisions for them?

AJ: Right, right. And that's a great question. Again, this is a very different customer in that respect. And there's a saying that I have in my team that I tell my team, “Be careful what you measure because that drives behaviors.” So, now that we are targeting a very specific set of customers, the scale is not the biggest priority on their mind at any given time, because they're building, they're prototyping, they're starting something small, and they're growing with this.

So, that's not something that they want to over optimize. It goes back to the point that I made: they want solutions that are more vanilla. But when you compare these concerns that you talked about, with the large cloud providers, they are real because their needs are very different. They're looking for a solution or Kubernetes clusters with more than thousand-plus nodes, and things along those lines. And to support that and to support those scenarios, yes, the complexity comes in by nature, and they have to build certain features and solutions to work around those limitations.

And that’s where things starts getting complicated because it's not a one size fits everyone scenario when you have these vast variety of customers coming in and trying to use the product that you build that was essentially built based on the feedback that you get—getting from your largest, and the biggest, and the highest paying customers. Goes back to the point that I made: be careful what you measure. Biggest thing that they measure based on the top cloud provider is the how many big multi-million-dollar deals are we getting? And those customers have very unique needs. Every customer will have a different feature requirements and sets, and if you rally around that, you end up building your product, by definition, that's complex.

Emily: When you're talking to your team about the conversations that you have with customers, what do you feel like is sort of hard to communicate? And I'm talking about when you're trying to translate, almost, these customer needs to your own technical team.

AJ: Yeah, I don't deal with—on a day-to-day—execution side of things, so I stay away from translating these requirements to my product team or engineering team. Instead, I end up introducing those guys to the customers and have them talk to them directly because nothing beats talking to the customers directly versus me getting the feedback and trying to relay the same to the broader team to go and build this product out. So, it's a big part of our DNA, it's a big part of our culture, on inviting customers, talking to them more frequently before we build anything. That works out. The learning that I get from talking to our customers is around defining the strategy and the vision for the company on who we want to do and who we want to be in two years, three years, or four years from now.

Emily: What do you see—for these smaller teams, smaller companies, what is the main motivation for using something like Kubernetes?

AJ: Yeah, it goes back to the point that I made: there's a huge amount of mindshare. So, you know what, you end up finding quite a bit of customers, they just want to use it because it's the buzzword, and then that's where they start. And the second motivation where the people who really, really know the technology, and know what it can do is to avoid the infrastructure management piece they had to deal with themselves around what happens when the—you know, once your machine goes down. How do I add certain things into [unintelligible]? How do I bring certain security isolation?

All those the orchestration piece that Kubernetes gives you by definition is very empowering to them to go and offload some of the manual work they used to do. The third tangent that I'm seeing, at least with DO and the customers that I speak to, that it’s a great platform and a cheap way for them to start learning about the technology because learning is a big part of what your customers are, right? So, as they learn about technology as the buzzword [unintelligible] Kubernetes keeps on getting hot, they come running into the documentation that we've created, they love that. And then they come in, spin up the clusters just to learn what this technology is. I've had some customers that are really, really large enterprises, but a smaller team within those enterprises, they're sending their dev teams to DO saying, “You know what? Great platform. Go learn about Kubernetes and see how the technology works because that's the easiest and fastest way, and the cheapest way to get that stuff done before you actually start using it.”

Emily: Do you think that most people come to learn about Kubernetes because Kubernetes is a buzzword?

AJ: Well, that's one part of it, right? There's one step, they’re trying to learn certain things, so you're seeing certain percent that they are just going for that mindshare they have. Then there is a set of customers who know what it is, and why they want to use it, and they are very thoughtful and mindful around what they're going to use that for. And that's where the SMB business comes in, or the more business side of the customer they come in, and to that point is where they just want to leverage and get rid of all the manual orchestration work they had to do with all these virtual machines that they had to be with. But then when you talk about the customers, you're talking about the customers who are dealing with those clusters of VMs. You’re not really talking about customers who's trying to just spin up a cluster with two or three nodes because that's the guy who's trying to learn.

Emily: And among your customers in particular, what do you think some of the misconceptions are, either about cloud-native in general or about Kubernetes specifically?

AJ: Yeah there's not really a misconception perspective. We… the conversations that I ended up having with them is not about the philosophy of the technology that has been evolved across us. I mean, obviously DO didn't invent the Kubernetes, or neither did Microsoft, so when you stay away from that conversation [unintelligible] that sort of conversation.

But the biggest misconception is that this is the way for me to build the cloud-native application because that's what the industry tells me to; that's what it looks like. So, some people go out and start their [unintelligible] applications on the top of Kubernetes thinking that's the only way; that's how I should do it, whether it really fits the bill or not. The reality is there might be and should be using some sort of a PaaS platform. They might be using some sort of a container solution that is a full-blown PaaS, and they don't really need the Kubernetes clusters. And that's where the misconception is, is because they just struggle making the right technology bets for the solution that they're trying to build and drive versus just catching on to the latest and greatest buzzword or the technologies that are being out there.

Emily: Why do you think this happens? Why do you think customers sometimes struggle to make the right technology decisions based on their actual business goals?

AJ: Because that's a really [laughs] hard problem to solve, at least when you're trying to build a business, you're not always ends up being the technology business. Technology does not all end up being your first [unintelligible], or the biggest skills set that you have. They’re just trying to pick and choose what is being influenced by the learning, or through forums and whatnot to go along with. There's hardly very few customers that I end up seeing, they make the technology decisions that are long term with the scalable solutions, and right coding language, and whatnot. It's mainly around what's the best and the latest that I can get my hands with? What I can learn quickly about? What are the resources are available? And let me start prototyping that fast. And soon.

Emily: What advice would you give to these businesses? Like, maybe they're not going to make the perfect technology decision, but if you have advice to maybe help them make a better technology decision, or just what to think about to try to improve that decision making process?

AJ: You know, I get that question asked quite a lot from people who are trying to start something new. And the right advice there is to focus on job to be done. What is the end goal that you're trying to solve for within a few weeks or a month, or maybe a quarter? What is the job to be done? And then look for various skills and comfort-level are versus trying choose for what is the latest and greatest and then spend a crazy amount of time trying to learn that technology because you don't have that skill set in-house or within you.

So, start with what you feel comfortable, where you have the skill sets, where you can make and prototype something quicker and sooner, but while keeping the core job to be done in the mind. And if the job to be done is not really technology-focused, it's solving some different business problem, then don't heavily pivot on picking the right technology.

Emily: Would you say this also sort of starts with an honest evaluation of what your… resources so to speak, what your skills are?

AJ: One hundred percent. That's the question you have to ask yourself as well. So, if you're comfortable with one coding language, go with it. Just don't go pick the latest and greatest coding language because that's the biggest buzzword, but then you end up finding yourself spending quite a bit of time learning that technology, but you're not really solving for the job that you [unintelligible].

Emily: What do you think are the best rationale that you see for using Kubernetes?

AJ: It's again, goes back to, in my opinion, Kubernetes is still infrastructure. People end up confusing that with the layer of platform as a service, and whatnot. The real rationale in using that technology is to optimize for some of the manual work you used to do around managing infrastructures and VM on a scale. SO some of the OS updates, the auto-upgrades, and what happens when the VM goes down? How do you add more VMs to your existing network and whatnot?

Those were very lengthy and time-consuming task, and all the orchestration pieces that Kubernetes provides, it's a great way to start and manage your existing infrastructure. I mean, that's the way I would recommend people to tiptoe into the technology, and that's where the majority of the customers are. Then there's a set of customers who are looking to build an high-level PaaS offering or a Software as a Service offering. These are more skilled people with the right talent and enough money to invest in the right place. Kubernetes is a great platform for them to build those sort of PasS and SaaS offering on top of that because it does not make any sense in trying to build those custom orchestration when you are building a solution that's going to serve tens of thousands of end-users and going to scale [unintelligible]. So, but those are pretty low percent of Kubernetes users, and that's where the industry trend seems to be, also, growing around people building larger businesses.

Emily: What advice would you have about making these longer-term technology decisions? When you first started talking about making good technology decisions, it was about what are we going to do this quarter? What about when you're thinking, “What are we going to do in the next two years? What are we going to do in the next five years?” How far out should you be thinking when you're making a technology decision?

AJ: Yeah, that's a pretty hard situation to be in. And not many will be trying to answer that question, especially the ones that I deal with, the customer segment that I'm talking about. On longer-term technology decisions, it's a very different set of customers who are building solutions for a larger customer segments and whatnot. The advice, again, remains the same. Start with something that you are comfortable with, keep an eye on what is the game plan is going to be when you start hitting a certain scale and keep an eye on your technical debt.

Again, you don't want to keep on piling on to technical debt for too long in trying to build a solution. So, whatever that you design, just make sure you have the capability, and you're choosing the technology that's going to be around and evolve when the time is right for you to go out and start paying your technical debt. The second aspect, I would ask them to invest in quite a bit of automation; investing quite a bit of work that it's not really sexy per se. When you're building features and writing code, building new features and designing the new product ideas is always fun, but then there's a workaround certain things on automation. There's always this one guy who does all these CI/CD automation, pipeline automation, making sure your code written a certain way, and those things; those are very critical investments.

If you do not make them a core part of your development process from get-go, you're going to have a really hard time in two years, or three years, or four years from now to keep up with what's going on and to keep up and automate all that code in the application that you wrote to scale with the business that is scaling.

Emily: What advice specifically do you have for startups? So, not just small companies; small companies that are looking to scale quickly. Do you think that there's any difference in how you recommend adopting Kubernetes or cloud-native versus, say, a small company that's not planning to scale massively in the next couple of years?

AJ: Yeah, I think the latter question is, is really even if you were to ask the founder of the company, she wouldn't be able to answer that, “We're not planning to scale that fast.” The scale comes; it surprises you. [laughs]. When it was about product-market fit, you're going to find the scale and that's going to happen. My recommendation is at a more generic level.

If you're starting something new, you’re building up from scratch, start cloud-native. Pick the right technology that are evolving, and start with open source because you're going to have the wave of innovations coming in. Why open source? They will outweigh trying to use something proprietary or trying to reinvent the wheel yourself. So, if you starting something new, start cloud-native, use open source. It allows you to quickly iterate and build things faster. And you benefit from a larger community contribution coming into those innovations. Like, at the end of the day, I like to say this, regardless of what proprietary solution you use or who you are and whatnot, open-source wins in the end.

Emily: Why do you think open source wins in the end?

AJ: It's because the larger community and multiple minds are working, trying to solve a similar problem. It’s far, far—much better. It's more inclusive. It's not as opinionated as when you would be when you're building your own proprietary software or trying to do something to solve your one specific problem. When you're building products and solutions that are going to be used across the world by different kinds of customers, your segments going to change, you really want to make a bet on a technology and a solutions that are built by different minds, more inclusive people, people with different opinion and thinking from yourself, and in the long run, that pays up.

The great example I can give you from my past life is I was fortunate to build a Platform as a Service for Azure. Back in the days, there was no Kubernetes. But what we ended up building to build that technology was just a really big custom orchestration exactly similar to what Kubernetes is today. The reality is fast forward to, you know, eight, nine years from now everybody's talking about Kubernetes, not that custom orchestration that we built. The business is working; that is great, but at the end of the day, what won was the open-source orchestration that allowed people to manage thousands and hundreds of thousands of VMs on a large scale, and that's winning.

Emily: Would you recommend open source even for small companies that, say, don't have a lot of expertise in—

AJ: 100 percent.

Emily: —A lot of people say open source is free like a puppy, right? So, you have to invest a lot in—

AJ: No, no, it's 100 percent because it's really easy to get started with open source. It's really easy to get something that's out there. There's a bigger community to try and help you out, and it gives immense pleasure to contribute back to something that you're consuming as well. I mean, this is, by definition, is a human nature. It keeps you more engaged and innovative when you're starting there. So, I think it’s when you're prototyping something, just always try and go with the open-source. You're going to get tons and tons of ideas and levers you could pull. If you’re going to run into some issues, somebody has, somewhere, ran into that and will have a solution. So, that's the easiest way to start.

Emily: Anything else that you'd like to add, that really sticks out about the business reasons that people are choosing Kubernetes or cloud-native?

AJ: Yeah. Again, like I said, the business reason, you know, majority of them… more than 50 percent are pick that to start something new, and learn in their perspective, and just use the benefits of the custom—or not the custom, but the open-source orchestration to deal with their few virtual machines, they were running before, or whatnot. Then there is a trend that is growing around the Kubernetes ecosystem where people are now building more Paas and SaaS solution on top of that, and these are the people who know how the technology works, who are the people who are contributing back to the Kubernetes ecosystem as well. But there are a layer of abstractions that have come in on top of Kubernetes where problems that were created by the technology—some of the innovations that you see from Google, Axure, that providing you a bunch of bells and whistles to go build those PaaS and Saas on top of Kubernetes. So, that's where—now the reason why people are using that.

Emily: All right, AJ, what is an engineering tool that you couldn't do your job without? Or maybe I should say, what's a tool that you just couldn't do your job without?

AJ: [laughs]. Well, I'm an engineer at heart, but I don't code anymore. I mean, it's been a while, but so if there's one tool that I can’t do my job without right now, into this world is definitely videoconferencing. When everybody's remote. And DO as a company was 70 percent remote before the pandemic. Now we’re 100 percent, so if that tool’s gone away, there is zero productivity. So, that's the biggest one that I have in my mind. And besides that, things like Slack. We work, live, breathe that, and it makes things pretty useful. So, those are the two things I can do my job without.

Emily: How can listeners connect with you or follow you?

AJ: Yeah, they can follow me on Twitter or connect with me on my LinkedIn. Both handles are my first name and J-O. It’s apurvajo, A-P-U-R-V-A-J-O. You know, I always love to hear from our customers, and a future prospect or anybody who wants to learn about the company and what we have is always exciting.

Emily: All right. Well, thank you so much for joining us.

AJ: Likewise. Thank you.

Emily: Thanks for listening. I hope you’ve learned just a little bit more about The Business of Cloud Native. If you’d like to connect with me or learn more about my positioning services, look me up on LinkedIn: I’m Emily Omier, that’s O-M-I-E-R, or visit my website which is emilyomier.com. Thank you, and until next time.

Announcer: This has been a HumblePod production. Stay humble.

View Details

The conversation covers:

  • Gou’s role as CTO of Portworx, and how he works with customers on a day to day basis.
  • Common pain points that Gou talks about with customers. Gou explains how he helps customers create agile and cost-effective application development and deployment environments.
  • The types of people that Gou talks to when approaching customers about cloud native discussions.
  • Why customers often struggle with infrastructure related problems during their cloud native journeys, and how Gou and his team help.
  • Common misconceptions that exist among customers when exploring cloud native solutions. For example, Gou mentions moving to Kubernetes for the sake of moving to Kubernetes.
  • Gou’s thoughts on state — including why there is no such thing as an end-to-end stateless architecture.
  • Some cloud native vertical trends that Gou is noticing taking place in the market.
  • The issue of vendor lock-in, and how data and state fit into lock-in discussions.
  • Gou’s opinion on where he sees the cloud native ecosystem heading.

Links

  • Portworx: https://portworx.com/
  • Portworx Blog: https://portworx.com/blog/
  • Gou Rao Email: mailto:gou@portworx.com

Transcript

Emily: Hi everyone. I’m Emily Omier, your host, and my day job is helping companies position themselves in the cloud-native ecosystem so that their product’s value is obvious to end-users. I started this podcast because organizations embark on the cloud naive journey for business reasons, but in general, the industry doesn’t talk about them. Instead, we talk a lot about technical reasons. I’m hoping that with this podcast, we focus more on the business goals and business motivations that lead organizations to adopt cloud-native and Kubernetes. I hope you’ll join me.

Emily: Welcome to The Business of Cloud Native, I'm your host Emily Omier, and today I am chatting with Gou Rao. Gou, I want to go ahead and have you introduce yourself. Where do you work? What do you do?

Gou: Sure. Hi, Emily, and hi to everybody that's listening in. Thanks for having me on this podcast. My name is Gou Rao. I'm the CTO at Portworx. Portworx is a leader in the cloud-native storage space. We help companies run mission-critical stateful applications in production in hybrid, multi-cloud, and cloud-native environments.

Emily: So, when you say you’re CTO, obviously that's a job title everyone, sort of, understands. But what does that mean you spend your day doing?

Gou: Yeah, it is an overloaded term. As a CTO, I think CTOs in different companies wear multiple hats doing different things. Here at Portworx, technically I'm in charge of this company strategy and technical direction. What does that mean in terms of my day to day activities? And it's spending a lot of time with customers understanding the problems that they're trying to solve, and then trying to build a pattern around what different people in different industries and companies are doing, and then identifying common problems and trying to bring solutions to market, by working with our engineering teams, that sort of address, holistically, the underlying areas that I see people try and craft solutions around, whether it's enabling an agile development environment for their internal developers, or cost optimization, there's usually some underlying theme, and my job is to identify what that is, and come up with a meaningful solution that addresses a wide segment of the market.

Emily: What are the most common pain points that you end up talking to customers about?

Gou: Over the past, I think, eight-plus years or so—I think the enterprise software space goes through iterations in the types of problems that are being solved. Over the past eight-plus years or so, it really has been around this—we use this term cloud-native—enabling cloud-native environments. And what does that really mean? In talking to customers, what this is really meant recently is enabling an agile application development and deployment environment. And let's even define what that is.

Me as an application developer, I have to rely on traditional IT techniques where there's a separate storage department, compute department, networking department, security department, and I have to interact with all of them just to develop and try out an application. But that really is impeding me as a developer from how fast I can iterate and build product and get it out there, so by and large, the common underlying theme has been, “Make that process better for me.” So, if I'm head of infrastructure how can I enable my developers to build and push product faster? So, getting that agility up in a sense where it makes—cost-wise, too, so it has to make cost sense—how do I enable an efficient, cost-efficient development platform? That has been the underlying theme. That sort of defines a set of technologies that we call cloud-native, and so orchestration tools like Kubernetes, and storage technologies like, hopefully, what we're doing at Portworx, these are all aimed at facilitating that. That's been sort of what we've been focused on over the past couple of years.

Emily: And when you talk to customers, do they tend to say, “Hey, we need to figure out a way to increase our development velocity?” Or do they tend to say, “We need a better solution for stateful applications?” What's the type of vocabulary that they're attempting to use to describe their problems, and how high-level do they usually go?

Gou: That's a good question. Both. So, the backdrop really is, “Increase my development velocity. Make it easier for me to put product out there faster.” Now, what does it take to get there? So, the second-order problems then become do I run in the public cloud, private cloud? Do I need help running stateful applications? So, these are all pillars that support the main theme here, which is increasing development velocity. So, the primary umbrella under which our customers are operating under is really around increasing the development velocity in a way that makes cost sense.

And if you double-click on that and look at the type of problems that they're solving, they would include, “How do I efficiently run my applications in a public cloud? Or a hybrid cloud? How do I enable workflows that need to span multiple clouds?” Again because maybe they're using cloud provider technologies, like either compute resources, or even services that a cloud provider may be offering, so that, again, all of this so that they can increase their development velocity.

Emily: And in the past, and to a certain extent now, storage was somewhat of a siloed area of expertise. When you're talking to customers, who are you talking to in an organization? I mean, is it somebody who's a storage specialist or is it someone who's not?

Gou: No, they're not. So, that's been one of the things that have really changed in this ecosystem, which is the shift away from this kind of like, hey, there's a storage admin and a storage architect, and then there's a compute admin or BM admin or a security admin, that's really not who are driving this because if you look at that—that world really thinks in terms of infrastructure first.

Actually, let me just take a step back for a second. One of the things that has actually changed here in the industry is this: a move from a machine-centric organization to an application-centric organization. So, let me explain what that means. Historically, enterprises have been run by data centers that have been run by a machine-centric control plane. This is your typical VMware type of control plane where the most important concept in a data center is a machine.

So, if you need storage, it's for a machine. If you need an IP address, it's for a machine. If you want to secure something, you're securing a machine. And if you look at it, that really is not what an enterprise or business is trying to solve. What enterprises and businesses are trying to do is put their product out there faster.

And so for them, what is more important? It's an application. And so what it's actually changed here? And one of the things that defines this cloud-native movement, at least in my mind, is this move from a machine-centric control plane to an application-centric control plane where the first-class citizen or the most important thing in an enterprise data center is not a machine, it's an application. And really, that's where technologies like Kubernetes and things like that come in.

So, now your question is, who do we talk to? We don't talk to a storage administrator or machine-centric administrator; we talk to the people that are more focused on building and harnessing that application-centric control plane. So, these are people like a cloud architect, or it's a CIO level—or a CTO level driven decision to enable their application developers to move faster. These are application owners, application architects, it's that kind of people. You had an additional question there which is, are they storage experts? And so by definition, these people are not. So, they know they need storage, they know they need to secure their applications, they know they need networking, but they're not experts in any one of those domains. There are more application-level architects and application experts.

Emily: Do you find that there tends to be some knowledge gaps? If so, is there anything that you keep sort of repeating over and over again when you have these conversations?

Gou: So, it's not about having a knowledge gap, it's more about solving an infrastructure problem—which is what storage is—is not necessarily their primary task. So, in a sense that they're not experienced with that, they know that they need infrastructure support, they know they need storage support, and networking support, but they expect that to be part of the core platform. So, one of the things that we've had to do—and I suspect others in this cloud-native ecosystem—is to take all of that the heavy lifting that storage software would do and then package it in such a way that it's easy to consume by application architects; easy to consume by people that are putting together this cloud-native platform. So, dealing with things like drive failures, or how do I properly [unintelligible] my data or RAID protect my data, or how do I deal with backups? These are things that they kind of know they need me to do, but they're not really experienced with it, and they expect the cloud-native software platforms to handle that functionality for them. So, in other words, doing the job of a storage admin, but in a form of software has been really important.

Emily: Do you find that there's any common misconceptions out there?

Gou: Yeah. With any new technology or evolving space, I think there are bound to be some misconceptions in how you approach things. So, just moving to Kubernetes for the sake of moving to Kubernetes, for instance, is a common—or improperly embracing is as a common mistake that we see. You really need to think about why you're moving to this platform, and if you're really doing it to embrace developer agility. For instance, one mistake we see, and this is especially true with the ecosystem we work in because Portworx is enabling storage technologies in Kubernetes. We see people try and take Kubernetes and leverage legacy infrastructure tools like either NFS or connecting their Kubernetes systems to storage arrays—which are really machine-centric storage arrays—over protocols like iSCSI, and you kind of look at that, and what is the problem with that?

And one way I like to describe it is if you take an agile platform like Kubernetes, and then you make that platform rely on legacy infrastructure, well, you're kind of bringing down the power of Kubernetes to the lowest common denominator here, which is your legacy platform. And your Kubernetes platform is only going to be as agile as the least nimble element in the stack. And so a mis-pattern that we see here is where people try and take them, and then they say, “Well, geez, I'm not getting the agility out of Kubernetes, I don't really see what all the fuss about this is. I'm still moving IT tickets around, and make developers still rely on storage admins.” And then we have to tell them, “Well, that's because you've kind of tied your Kubernetes platform down to legacy infrastructure, and let's now think about modernizing that infrastructure.” And that’s sort of—I do find myself in a spot where we have to educate the partners about that. And they eventually see that and hopefully, that's why they work with us.

Emily: Why do you think they tend to make that mistake?

Gou: It's what’s lying around, right? It's the ease of just leveraging the tools and the equipment you already have, not really fully understanding the problem all the way through. It takes time for people to try out a pattern—take Kubernetes, try and see what you have lying around, connect it to learn from mistakes. And so that really takes some time. They look at others to see what others are doing, and it's not immediately evident. I think, with any new technology, there's always some learning that goes along with it.

Emily: Well, let's talk a little bit about stateful apps in general. Why do people need stateful apps? What business problem do stateful apps deal with, or make it possible to solve that you just can't do with stateless?

Gou: Sure. Yeah very rarely is there really something that is a stateless app, so unless you're doing some sort of ephemeral experiment—and there are certain use cases for that—especially in enterprises, you always rely on data. And data is essentially your state in some shape or form, whether it's short-lived state, long-lived state, there's always some state and data that's involved in an application. And people try to think of things in terms of stateless, and what that really means is, they're punting the state problem to some other part of their overall solution; it doesn't really go away.

What do I mean by that? Well, if you put all your stateless apps on your cloud-native platform, where did the state go? You're probably dealing with it in some VM or some bare-metal machine. It didn't really go away; it's there, and you're still left with trying to connect to it, and access it, and manage it. And then you start wondering, what kind of northbound impact is that having on your, what you think is your stateless side of things.

And it has an impact. Anytime you need to make changes to your data structure, it's not like your stateless side of things is not impacted, so it kind of halts the entire pipeline. So, what we see happening as a pattern is people finally understanding that there's really no such thing as an end-to-end stateless architecture—especially in enterprises—and that they need to embrace the state and manage it in a cloud-native way. And so, that's really where a lot of this talk you see around these days, around how do you manage stateful applications in Kubernetes, that's where that is coming from. How do you govern it? How do you enforce things like RBAC on it? How do you manage its accessibility? How do you deal with its portability, because state has gravity? These are the main topics these days that people have to think about.

Emily: Can you think of any misconceptions other than the ones that you've just mentioned that are specifically related to state?

Gou: Sure. I think costs—well, less a misconception than I think not fully appreciating or understanding the impact that this would have, and it really is around cost and especially when running in the public cloud. People underestimate the cost associated with state in the public cloud, and especially when you're enabling an agile platform, with agility comes—really determines automation, and people tend to do a lot of things and so, when you have a lot of state laying around, the cost can add up. Simple example is if you have a lot of volumes that are sitting around in, let's say, Amazon, well, you're paying for those bills.

And density is another related concept, which is, if you're running thousands of applications, and thousands of databases, or thousands of volumes, how do you manage that with not spending that much money on the resources? How do you virtually manage that? So, cost planning in the public cloud is something that I see is so commonly overlooked or not really fully thought through. And it does come to bite people down the line because those bills stack up quickly. So, coming up with an architecture where that is well thought through becomes really important.

Emily: When do people tend to think about state when they're thinking about a digital transformation towards cloud-native?

Gou: I think, Emily, state is something that is front and center of any enterprise architecture because it really—look at it this way, any application you're going to build, generally speaking, revolves around the data structures and the data it operates on. You really typically start with—and there's a few different ways in which you start with cracking your application stack, right? One is you start with the user experience, like what end-user experience you want to fulfill, what is the problem you're solving? Another core element is then what are the data structures needed to solve that problem? So, state becomes an important thing right from day one.

Now, in terms of managing it, people can choose to defer that problem saying, “I’m not going to deal with managing the state in my cloud-native platform.” This goes to your comment about stateless architectures. Again, if they do that, then they're punting the problem until the point where they actually need to implement it and manage it in production. Either way, that problem comes up. You're thinking about the data structures and development time for sure, you can avoid that. In terms of embracing how you're going to run it, it definitely comes into place as you're rolling up your application as you're getting close to production because somebody has to manage that.

Emily: And there are probably people out there who would argue, hey, cloud-native means stateless. What would you say to those people?

Gou: Yeah, I think we've had this debate many years ago, too. This notion of ‘twelve-factor applications’ is kind of where it started, and I think people realized that, unless you're dealing with an architecture where state is sort of a small subset of the overall product offering, where really a lot of value is in maybe your UI, or it's a web app where people are interacting with each other and there's just a small amount of data sitting in a database that's recording a few messages, there really is no such thing as a safe plus architecture. So, in that case, what you could do is you architect your database, maybe just once and you’re really not touching it—any changes you’re making to your application are more cosmetic—then you can put your state somewhere else and an external database, have an endpoint that your applications interact with. Keep in mind, though, that still—that's not stateless; you've just put your state outside of your cloud-native platform.

But what I would say to your question, what would you say to somebody that says that's the way to do things, my point is that's not how enterprise applications and architectures work. They're more complex than that, and the state is more central, the data is more central to the application, where changes intimately involve changing how your data structure is organized, changing tables around. In that case, it doesn't make sense to treat that outside of your cloud-native stack: because you're making so many changes to it, you're then going to lose the agility that the cloud-native platform can offer compared to bringing those microservice databases into your Kubernetes platform, letting the developers make the data structure changes they need to do as frequently as they need to, all within the same platform, within the same control plane. That makes a better enterprise architecture.

Emily: Sort of a different type of question. But I'm just curious if there's anything that continues to surprise you about the cloud-native ecosystem, even about your conversations with customers, what continues to surprise you, but also what continues to surprise them?

Gou: I was a little bit more surprised a couple of years ago, but does continue to surprise me is the mixing of the cloud-native and the legacy tools that still happens, and it's not just with storage, I see the same thing happening around security, or even networking. And some of that, not so much as surprise as and it makes less sense to me—and eventually I think people realize it—that they make this change, and then it sort of looks like a lateral move to them because they didn't get the agility they wanted. If somebody has to roll out an application and they have to sit through a machine type of security audit, for instance, if they're doing security, and they're not leveraging the new tools that do container-native security, those kinds of things do raise a flag and trying to work with our partners and customers and trying to point these things out, I still do see some of that happening. And I think it has been getting better over time because people learn from their peers within their industries, whether it's banking—you'll see banks kind of look at each other, and they develop similar architectures. It's a close-knit ecosystem, for instance.

Emily: Actually, do you see any specific trends that are vertical-specific?

Gou: So, industry-specific, right? So, for instance, in the financial sector, they're certainly trends, whether it's embracing hybrid-cloud technologies, and they kind of do things similarly. An example is, some industries are okay with completely relying on one single cloud provider. Generally speaking—and I'm just giving an example in the financial space—we've seen that not be the case, where they kind of want more of a multi-cloud or cloud-neutral platform, they probably will run on-prem and in the public cloud, but they don't want to be locked into a particular cloud provider. And so that's, for example, a trend in that industry. So, yeah, different industries have different kinds of specific features that they're looking for and architectures that they're circling around.

Emily: You brought up lock-in. Lock-in is a big topic with a lot of people, but often the idea of data gravity doesn't make its way into the primary conversations about lock-in. How does data and state fit into these lock-in discussions?

Gou: That's a really good point, and I'll break down into two things. So, data has gravity, and ultimately if you have petabytes of data sitting around somewhere, it becomes harder and harder to move that, but even before that, one of the most important things that we try and teach people, and I think people, kind of, realize this themselves is lock-in starts even before the data gravity has become an issue. It starts with, is your enterprise architecture or platform itself, just completely relying on a particular provider’s APIs and offerings? And that's where we really caution our customers and partners to say, that's the layer at which you want to build your—that’s your demarcation point, which is run your services in a cloud provider but don't lock your architecture around relying on that cloud provider providing those services.

So, instance as database, if you're running a database, it's better to run your database on your platform in the public cloud, as opposed to putting all of your data in the cloud providers database because then you're really locked in, not just by way of data, but even by way of your enterprise architecture. So, that's one of the first things that I think people are looking for which is, how do I build this cloud-agnostic platform? Forget cloud-native, but cloud-agnostic platform where I can run all of my databases at ease, but I've not locked in my database to a cloud provider. Once you do that, then you can start breaking up your applica—especially in enterprises, it's not like you just have one very large database; you typically have many, many, many small databases, so it becomes easy to start with portability, and you can have sets of your—if you need to—applications run in different cloud providers and your ways to connect these and this is the notion of hybrid-cloud, which is becoming real.

Emily: What do you see as the future? Where do you see this ecosystem going?

Gou: You know, there's a lot happening in different segments, right. And 5G, for example, is a huge enabler of Kubernetes, or consumer of cloud-native technologies, and there's just a lot happening over there. Just around edge to core compute, and patterns that are emerging with how you move data from the edge to the core or vice versa, and how you distribute data at the edge. So, there's a lot of innovations happening there in the 5G space. Every segment has something interesting going on. From Portworx’s standpoint, what we're doing over the next couple of years and helping people in two main areas.

I mentioned 5G, so enabling workflows that are truly cloud-native. What do I mean by that? Driven by Kubernetes, that allow the developers to create higher-level workflows where they can either move data between Kubernetes clusters—again, it doesn't have to be between cloud providers; just to take a step back for a second. Whenever somebody is running a cloud-native platform, what we find is that it's not like an enterprise has one very large Kubernetes clusters. Typically people are managing fleets of Kubernetes clusters, so whether they're doing blue/green deployments, or they just have different clusters for different types of applications, or they're compartmentalized for different departments, we find that people having to move and facilitate movement of applications between clusters is very important. So, that's an area where we're focusing on; certainly makes sense in the 5G space.

The other area is around putting in more AI into the platform. So, here's an example. Does every enterprise out there, does every developer need to be—if they're running stateful applications, do they need to become a database expert to run it? And our answer is no. We want to bring in this notion of self-driving technologies into stateful applications where—hopefully with the Portworx software—if you're running stateful applications, the Portworx software is managing the performance of the database, maybe even the sharding of the database, the vertical and horizontal scaling of it, monitoring the database for performance problems, hotspots, reacting to it.

We've introduced some technologies like that already over the past couple of years. An example is capacity management. What we’ve found is—and I think you asked this question early on—people that are running stateful applications on their Kubernetes platform, they're not storage experts. People routinely make mistakes when it comes to capacity planning. One of the things that we found is people had underestimated how much space they're going to use or how much storage they're going to use, and so they ran out of capacity in their cloud-native platform. Why should it be their problem to deal with that? So, we added software that's intelligent enough to detect these kind of cases and react to it. Similarly, we'll do the same thing around vertical scaling, reducing hotspots. And bringing in that kind of intelligence and little platform is something not just Portworx but we expect other people in this ecosystem to solve.

Emily: Do you think there's any business problems that customers talk about that you don't have a good answer to? Or that you don't think the ecosystem has a good answer to, yet?

Gou: Yeah, no, it’s a good question. So, making the platform simple. So, the simplicity is one aspect. I think we do see enterprises struggling with the cloud-native space being slightly complex enough with so many technologies out there. Kubernetes itself, certainly, it has its own learning curve.

So, making some of those things more simple and cookie-cutter so that enterprises can simply focus on just their application development, that is an area that needs to still be worked on. And we see enterprises, I wouldn't say really struggling, but trying to wrap their head around how to make the platform more simple for their developers. There are projects in the cloud-native space that are focused on this. I think a lot of this is going to come down to establishing a set of patterns for different use-cases. The more material and examples that are out there and use-cases that are out there will certainly help.

Emily: Do you have an engineering tool that you can't live without? If so, what is it?

Gou: [laughs]. My go-to is obviously GDB. It's our debugger environment. I certainly wouldn't be able to develop without that. When I am looking into insights into how a certain customer’s environment is doing, Prometheus and Grafana are, sort of, my go-to tools in terms of looking at metrics, and application performance health, and things like that.

Emily: How could listeners connect with you or follow you?

Gou: On our site portworx.com, P-O-R-T-W-O-R-X-dot-com. We frequently blog on that site; that would be a good way to follow what we're up to and what we're thinking. I certainly respond to people via email. So, reach out to me if you have any questions. I’m pretty good replying to that: Gou—that’s G-O-U—at portworx.com. I don't nearly tweet as much as our PR team would like me to tweet so there's not too much information there, unfortunately.

Emily: Thank you so much for joining us on The Business of Cloud Native.

Gou: Thank you, Emily. Thanks for having me.

Emily: Thanks for listening. I hope you’ve learned just a little bit more about The Business of Cloud Native. If you’d like to connect with me or learn more about my positioning services, look me up on LinkedIn: I’m Emily Omier, that’s O-M-I-E-R, or visit my website which is emilyomier.com. Thank you, and until next time.

Announcer: This has been a HumblePod production. Stay humble.

View Details

The conversation covers:

  • Why Kevin helped launch Single Music, where he currently provides SRE and architect duties.
  • Single Music’s technical evolution from Docker Swarm to Kubernetes, and the key reasons that drove Kevin and his team to make the leap.
  • What’s changed at Single Music since migrating to Kubernetes, and how Kubernetes is opening new doors for the company — increasing stability, and making life easier for developers.
  • How Kubernetes allows Single Music to grow and pivot when needed, and introduce new features and products without spending a large amount of time on backend configurations.
  • How the COVID-19 pandemic has impacted music sales.
  • Single Music’s new plugin system, which empowers their users to create their own middleware.
  • Kevin’s current project, which is a series of how-to manuals and guides for users of Kubernetes.
  • Some common misconceptions about Kubernetes.

Links

  • Single Music
  • Traefik Labs
  • Twitter: https://twitter.com/notsureifkevin?lang=en
  • Connect with Kevin on LinkedIn: https://www.linkedin.com/in/notsureifkevin

Emily: Hi everyone. I’m Emily Omier, your host, and my day job is helping companies position themselves in the cloud-native ecosystem so that their product’s value is obvious to end-users. I started this podcast because organizations embark on the cloud naive journey for business reasons, but in general, the industry doesn’t talk about them. Instead, we talk a lot about technical reasons. I’m hoping that with this podcast, we focus more on the business goals and business motivations that lead organizations to adopt cloud-native and Kubernetes. I hope you’ll join me.

Emily: Welcome to The Business of Cloud Native. I'm Emily Omier, your host, and today I am chatting with Kevin Crawley. And Kevin actually has two jobs that we're going to talk about. Kevin, can you sort of introduce yourself and what your two roles are?

Kevin: First, thank you for inviting me on to the show Emily. I appreciate the opportunity to talk a little bit about both my roles because I certainly enjoy doing both jobs. I don't necessarily enjoy the amount of work it gives me, but it also allows me to explore the technical aspects of cloud-native, as well as the business and marketing aspects of it. So, as you mentioned, my name is Kevin Crawley. I work at a company called Containous. They are the company who created Traefik, the cloud-native load balancer.

We've also created a couple other projects, and I'll talk a little bit about those later. For Containous, I'm a developer advocate. I work both with the marketing team and the engineering team. But also I moonlight as a co-founder and a co-owner of Single Music. And there, I fulfill mostly SRE type duties and also architect duties where a lot of times people will ask me feedback, and I'll happily share my opinion. And Single Music is actually based out of Nashville, Tennessee, where I live, and I started that with a couple friends here.

Emily: Tell me actually a little bit more about why you started Single Music. And what do you do exactly?

Kevin: Yeah, absolutely. So, the company started out of really an idea that labels and artists—and these are musicians if you didn't pick up on the name Single Music—we saw an opportunity for those labels and artists to sell their merchandise through a platform called Shopify to have advanced tools around selling music alongside that merchandise. And at the time, which was in 2016, there weren't any tools really to allow independent artists and smaller labels to upload their music to the web and sell it in a way in which could be reported to the Billboard charts, as well as for them to keep their profits. At the time, there was really only Apple Music, or iTunes. And iTunes keeps a significant portion of an artist's revenue, as well as they don't release those funds right away; it takes months for artists to get that money.

And we saw an opportunity to make that turnaround time immediate so that the artists would get that revenue almost instantaneously. And also we saw an opportunity to be more affordable as well. So, initially, we offered that Shopify integration—and they call those applications—and that would allow those store owners to distribute that music digitally and have those sales reported in Nielsen SoundScan, and that drives the Billboard Top 100. Now since then, we've expanded quite considerably since the launch.

We now report on sales for physical merchandise as well. Things like cassette tapes, and vinyl, so records. And you'd be surprised at how many people actually still buy cassette tapes. I don't know what they're doing with them, but they still do. And we're also moving into the live streaming business now, with all the COVID stuff going on, and there's been some pretty cool events that we've been a part of since we started doing that, and bands have gotten really elaborate with their live production setups and live streaming.

To answer the second part of your question, what I do for them, as I mentioned, I mostly serve as an advisor, which is pretty cool because the CTO and the developers on staff, I think there's four or five developers now working on the team, they manage most of the day-to-day operations of the platform, and we have, like, over 150 Kubernetes pods running on an EKS cluster that has roughly, I'd say, 80 cores and 76 gigabytes of RAM. That is around, I'd say about 90 or 100 different services that are running at any given time, and that's across two or three environments, just depending on what we're doing at the time.

Emily: Can you tell me a little bit about the sort of technical evolution at Single? Did you start in 2016 on Kubernetes? That's, I suppose, not impossible.

Kevin: It's not impossible, and it's something we had considered at the time. But really, in 2016, Kubernetes, I don't even think there wasn't even a managed offering of Kubernetes outside of Google at that time, I believe, and it was still pretty early on in development. If you wanted to run Kubernetes, you were probably going to operate it on-premise, and that just seemed like way too high of a technical burden. At the time, it was just myself and the CTO, the lead developer on the project, and also the marketing or business person who was also part of the company. And at that time, it was just deemed—it was definitely going to solve the problems that we were anticipating having, which was scaling and building that microservice application environment, but at the time, it was impractical for myself to manage Kubernetes on top of managing all the stuff that Taylor, the CTO, had to build to actually make this product a reality.

So, initially, we launched on Docker Swarm in my garage, on a Dell R815, which was like a, I think was 64 cores and 256 gigs of RAM, which was, like, overkill, but it was also, I think it cost me, like, $600. I bought it off of Craigslist from somebody here in the area. But it served really well as a server for us to grow into, and it was, for the most part, other than electricity and the internet connection into my house, it was free. And that was really appealing to us because we really didn't have any money. This was truly a grassroots effort that we were just—we believed in the business and we thought we could quickly ramp up to move into the Cloud.

So, that's exactly what happened though. Like, we started making money—also, this was never my full-time job. I started traveling a lot for my other developer relations role. I worked at Instana before Containous. Eventually, the whole GarageOps thing just wasn't stable for the business anymore.

I remember one time, I think I was in Scotland or somewhere, and it was, like, two o'clock in the morning at home here in Nashville, and the power went out. And I have a battery backup, but the power went out long enough to where the server shut down, and then it wouldn't start back up. And I literally had to call my wife at two o'clock in the morning and walk her through getting that server back up and running. And at that point in time, we had revenue, we had money coming in and I told Taylor and Tommy that, “Hey, we're moving this to AWS when I get back.” So, at that point, we moved into AWS. We just kind of transplanted the virtual machines that were running Docker Swarm into AWS. And that worked for a while, but up until earlier this year, it became really apparent that we needed to switch the platform to something that was going to serve us over the next five years.

Emily: First of all, is ‘GarageOps’ a technical term?

Kevin: I mean, I just made it up.

Emily: I love it.

Kevin: I mean, it was just one of those things where we thought it was a really good idea at the time, and it worked pretty well because, in reality, everything that we did, up into that point was all webhook-based, it was really technically simple. But anything that required a lot of bandwidth like the music itself, it went directly into AWS into their S3 buckets, and it was served from there as well. So, there wasn't really any of this huge bandwidth constraint that we had to think about, that ran in our application itself. It was just a matter of really lightweight JSON REST API calls that you could serve from a residential internet connection if you understand how to set all that stuff up. And at the time, I mean, we were using Traefik, which version 1.0 at the time, and it worked really well for getting all this set up and getting it all working, and we leveraged that heavily.

And at that time in 2016, there wasn't any competitor to Traefik. You would use HAProxy or you use NGINX, and both of those required a lot of hand-holding, and a lot of configuration, and it was all manual, and it was a nightmare. And one of the cool things about Docker Swarm and Traefik is that once I had all the tooling set up, it all sort of just ran itself. And the developers, I don't know around 2017 or ’18, we had hired another developer on the staff. And realistically, if they wanted to define a new service, they didn't have to talk to me at all.

All they did was create a new repo in GitHub, change some configuration files in the tooling we had built—or that I had built—and then they would push their code to GitLab, and all the automation would just take over and deploy their new service, and it would become exposed on the internet, if it was that type of a service, it was an API. And it would all get routed automatically. And it was really, really nice for me because I really was just there in case of the power went out in my garage, essentially.

Emily: You said that up until earlier this year, this was more or less working, and then earlier this year, you really decided it wasn't working anymore. What exactly wasn't working?

Kevin: There were a few different things that led us to switching, and the main one was it seemed like that every six to twelve months, the database backend on the Swarm cluster would fall over. For whatever reason, it would just—services would stop deploying, the whole cluster would seemingly lock up. It would still work, but you just couldn't deploy or change anything, and there was really no way to fix it because of how complicated and how I want to say how complex the actual databases and the data that's been stored in it because it's mostly just stateful records of all the changes that you've made to the cluster up until that point. And there was no real easy way to fix that other than just completely tearing everything down and building it up from scratch. And with all the security certificates, and the configuration that was required for that to work, it would literally take me anywhere between five to ten hours to tear everything apart, tear everything down, set up the worker nodes again, and get everything reestablished so that we could deploy services again, and the system was accepting webhooks from Shopify, and that was just way too long. Earlier this year, actually we crossed into, I want to say in January, we had over 1400 merchants in Shopify sending us thousands of orders every day, and it just wasn't acceptable for us to have that length of downtime 15, 20, 35 minutes, that's fine but several hours just wasn't going to work.

Our reputation up until that point had been fairly solid. That issue or incident hadn't happened in the past eight months, but we were noticing some performance issues in the cluster, and in some cases where we were having to redeploy services two, three times for those services to apply, and that was sort of like a leading indicator that something was going to go wrong pretty soon. And it was just a situation where it was like, “Well, if we're going to have to go offline anyways, let's just do the migration.” And it just so happened that in April, I was laid off from my job at Instana and I was fortunate enough to be able to find a new job in, like, a week, but I knew that I wanted to complete this migration, so I went ahead and decided to put off starting the new job for a month. And that gave me the means, and the opportunity and the motive to actually complete this migration.

There were some other factors that played into this as well, and that included the fact that in order to get Swarm stood up in 2016, I had to build a lot of bespoke tooling for the developers and for our CI/CD system to manage these services in the staging and production environment, handling things like promotion and also handling things like understanding what versions of the services are running in the cluster at any given time, and these are all tools that are widely available today in Kubernetes. Things like K9s, or Lens, or Helm, Kustomize, Skaffold, these are all tools that I essentially had to build myself in 2016 just to support a microservice environment, and it didn't make sense for us to continue maintaining that tooling and having to deal with some of their limitations because I didn't have time to keep that tooling fresh and keep it up-to-date and competitive with what's in the landscape today, which are the tools that I just described. So, it just made so much sense to get rid of all that stuff and replace it with the tools that are available today by the community and has infinitely more resources poured into them than I was ever able to provide, or I will ever be able to provide even as a single person working on a project. The one that was sort of lingering in the background was the fact that we have here recently started doing album releases, and artists are coming to us where they will sell hundreds of thousands of albums within a very short period of time, within several hours, and we were reaching the constraints of some of our database and our backend systems to where we needed to scale those horizontally. We had, kind of, reached the vertical limits of some of them, and we knew that Kubernetes was going to give us these capabilities through the modern operator pattern, and through just the stateful tooling that has matured in Kubernetes that wasn't even there in 2016, and wasn't something that we could consider, but we can now because the ecosystem has matured so much.

Emily: So, yeah, it sounds like basically you were running up against some technical problems that were on the verge of becoming major business problems: the risk of downtime, and the performance issues, and then it also sounds like some of the technical architecture was limiting the types of products, the types of services that you could have. Does that sound about right?

Kevin: Yeah, that's a pretty good summary of it. I think that one of the other things that we had to consider too was that the Single ecosystem, like the Single Music line of products has become so wide and so vast—I think we're coming up on five or six different product lines now—and developers need an 8 core laptop with 32 gigs of RAM just to stand up our stack because we're starting to use things like Kafka and Postgres to do analytics on all this stuff, and we're probably going to get to the point within the next 18 months to where we can't even stand up the full Single Music stack on a local machine. We're going to have to leverage Kubernetes in the Cloud for developers to even build additional products into the platform. And that's just not possible with Swarm, but it is with Kubernetes.

Emily: Tell me a little bit about what has changed since making the migration to Kubernetes. And I'm actually also curious, the timeframe when this happened is really interesting, and you talked a little bit about offering these streaming services for musicians. I mean, it's an interesting time to be in the music industry. Interesting, probably in both the exciting sense and also negative sense. But how have things changed? And how has Kubernetes made things possible that maybe wouldn't have been possible otherwise?

Kevin: I think right now, we're still on the precipice, or on the leading edge of really realizing the capabilities that Kubernetes has unlocked for the business. I think right now, I mean, the main benefit of it has been just a overwhelming sense of comfort and ease that has been instilled into our business side of the company, our executive side, if you will. The marketing and—of course, the sales and marketing people don't really know that much about the technical challenges that the engineering side has, and what kind of risk we were at when we were using Swarm at the time, but the owner did. There's three co-owners of the company, it's myself, Taylor, and Tommy. And Taylor, of course, is the CTO, and he is very well have the risk because he is deeply invested in the platform and understands how everything works.

Now, Tommy, on the other hand, he just cares, “Is it up?” Are customers getting what their orders—are they getting their music delivered? And so, right now it's just there's a lot more confidence in the platform behaving and operating like it should. And that's a big relief for the engineers working on the project because they don't have to worry about whether or not the latest version of their service that they deployed has actually been deployed; or if the next time they deploy, are they going to bring down the entire infrastructure because the Swarm database corrupts, or because the Swarm network doesn't communicate correctly like it missed routes. We had issues where staging versions of our application would answer East-West traffic—like East-West request traffic that is supposed to go in between the services that are running in the cluster—like staging instances would answer requests that were coming from production instances when they weren't supposed to. And it's really hard to troubleshoot those problems, and it's really hard to resolve those. And so right now it's just a matter of stability.

The other thing that is enabling us to do is handle the often difficult task of managing database migrations, as well as topic migrations, and, really, one-off type jobs that would happen every once in a while just depending on new products being introduced or new functionality to existing products being introduced. And these would require things like migrations in the data schema. And this used to have to be baked into the application itself, and this was really sometimes kind of tricky to manage when you start talking about applications that have multiple replicas, but with Kubernetes, you can do things like tasks, and jobs, and things that are more suited towards these one-off type activities that you don't have to worry about a bunch of services running into each other and stepping on each other's feet anymore. So, this, again, just gives a lot of comfort and peace of mind to developers who have to work on this stuff. And it also gives me peace of mind because I know ultimately, that this stuff is just going to work as long as they follow the best practices of deploying a Kubernetes manifest and Kubernetes objects, and so I don't have to worry about them breaking things per se, in a way in which they aren't able to troubleshoot, diagnose, and ultimately fix themselves.

So, it just creates less maintenance overhead for me because as I mentioned at the beginning of the call, I don't get paid by Single Music, unless of course, they go public or they sell. But I'm not actually a full-time employee. I'm paid by Containous, that's my full-time job, so anything that allows me to have that security and have less maintenance work on my weekends is hugely beneficial to my well being and my peace of mind, as well. Now, the other part of the question you had, as well, is in terms of how are we transitioning, and how are we handling the ever-changing landscape of the business? I think one of the things that Kubernetes lets us do really well is pivot and introduce these new ideas and these new concepts, and these new services to the world.

We get to release new features and products all the time because we're not spending a ton of time having to figure out, “Well, how do I spin up a new VM, and how do I configure the load balancer to work, and how do I configure a new schema in the database?” The stuff, it's all there for us already to use, and that's the beauty of the whole cloud-native ecosystem is that all these problems have been solved and packaged in a nice little bundle for us to just scoop up, and that enables our business to innovate and move fast. I mean, we try not to break things, but we do. But for the most part, we are just empowered to deliver value to our customers.

And for instance the whole live-streaming thing, we launched that over the course of, maybe, a week. It took us a week to build that product and build that capability, and of course, we've had to invest more time into it as time has gone on because not only do our customers see value in it, we see value in it, and we see value in investing additional engineering and business marketing hours into selling that product. And so again, it's just a matter of what Kubernetes, and the cloud-native ecosystem in general—and this includes Swarm to some extent because we could not have gotten to where we did without Swarm in the beginning, and I want to give it its proper dues because, for the most part, it worked really well, and it served our needs, but it got to the point where we kind of outgrew it, and we wanted to offload the managing of our orchestrator to somebody else. We didn't want to have to manage it anymore. And Kubernetes gave us that.

Emily: It sounds like, particularly when we're talking about the live streaming product, that you were able to build something really quickly that not only helped Single’s business but then obviously also helped a lot of musicians, I'm assuming at least. So, this was a way to not just help your own business, but also help your customers successfully pivot in a time of fairly large upheaval for their industry.

Kevin: Right. And I think one of the cool things that we experienced through the pandemic is that we saw a fairly sharp rise in sales in general in music, and I think it kind of speaks to the human nature. And what I mean by that, is that music is something that comforts people and gives people hope, and also it's an outlet. It's a way for people to, I don't want to say, disconnect because that's not really what I mean, but it gives them a means to experience something outside of themselves. And so it wasn't really that big of a surprise for us to see our numbers increase.

And, I mean, the only thing that kind of did surprise—I mean, it's not a surprise now in retrospect, but one of the things that we observed as well, as soon as all the George Floyd protests started happening across the United States, the numbers conversely dropped, and at that point, we realized that there was something more important going on in the world. And we expected that and we were… it was just an interesting observation for us. And right now, I mean, we're still seeing growth, we're still seeing more artists and more bands coming online, trying to find new ways to innovate and to try to sell their music and their artwork, and we love being a part of that, so we're super stoked about it.

Emily: That actually might be a good spot for us to wrap up, but I always like to give guests the opportunity to just say anything that they feel like has gone unsaid.

Kevin: Well, I mean, one of the things I do want to talk about a little bit is some of the stuff that we're doing at Containous as well. As a developer advocate, I think one of the things that I really enjoy in that aspect is that this gives me an opportunity to work closely with engineers in a way in which—a lot of times, they don't have an opportunity to experience the marketing and the business side of the product, and the fact that I can interact with my community and I can work with our open-source contributors and help the engineers realize the value of that is incredible. A few things that I've done at Containous since I've joined is we are working really hard at improving our documentation and improving the way in which developers and engineers consume the Traefik product. We also are working on a service mesh, which is a really cool way for services to talk to each other. But one of the things that we've recently launched two that I want to touch on is our plugin system, which is a fairly highly requested feature in Traefik.

And we launched it with Pilot, which is a new product that allows the users of Traefik to install these plugins that manipulate the request before it gets sent to the service. And that means our end-users are now empowered to create their own middleware, in essence. They're able to create their own plugins. And this allows them really unlimited flexibility in how they use the Traefik load balancer and proxy. The other thing that we're working on, too, is improving support for Kubernetes.

One of the surprises that I had when migrating from Traefik version 1 to Traefik 2, when we did the Single migration to Kubernetes, was once I figured out the version two configuration, it was really easy to make that migration, but it was difficult at first to make the translation between the version 1 schema of the configuration into the version 2. So, what we're working on and what I'm working on right now with our technical writer, is a series of how-tos and guides for users of Kubernetes to be empowered in the same way that we are at Single Music to quickly and easily manage and deploy their microservices across their cluster. With that, though, I mean, I do want to talk one more thing, on maybe some misconceptions about cloud-native and Kubernetes.

Emily: Oh, yes, go ahead.

Kevin: Yeah, I mean, I think one of the things that I hear a lot of is that Kubernetes is really hard; it's complex. And at first, it can seem that way; I don't want to dispute that, and I don't want to dismiss or minify people's experience. But once those basic concepts are out of the way, I think Kubernetes is probably one of the easiest platforms I've ever used in terms of managing the deployment and the lifecycle of applications and web services. And I think probably the biggest challenge is for organizations and for engineers who are trying to adopt Kubernetes is that in some ways, perhaps they're trying to make Kubernetes work for applications and services that weren't designed from the ground up to work in a cloud-native ecosystem. And that was one of the things that we had the advantage of in 2016 was even though we were using Docker Swarm, we still followed something which was called the ‘Twelve-Factor App’ principle.

And those principles really just laid us out for a course of smooth uninterrupted, turbulence-free flying. And it's been really an amazing journey because of how simple and easy that transition from Docker Swarm into Kubernetes was, but if we had built things the old way, using maybe Packer and AMIs and not really following the microservice route, and hard coding a bunch of database URLs and keys and all kinds of things throughout our application, it would have been a nightmare. So, I want to say to anybody who is looking at adopting Kubernetes, and if it looks extremely daunting and technically challenging, it may be worth stepping back and looking at what you're trying to do with Kubernetes and what you're trying to put into it, and if there needs to be some reconciliation at what you're trying to do with it before you actually go forth and use something like Kubernetes, or containers, or this whole ecosystem for that matter.

Emily: Let me go ahead and ask you my last question that I ask everybody which is, do you have a software engineering tool that you cannot live without, that you cannot do your job without? If so, what is it?

Kevin: Yeah, I mean, Google’s probably… [laughs] seriously, it's one of my most widely used tools as a developer, or as a software engineer, but in terms of, like, it really depends on the context of what I'm working in. If I'm working on Single Music, I would have to say the most widely used tool that I use for that is Datadog Because we have all of our telemetry going to there. And Datadog gives me a very fast and rapid understanding of the entire environment because we have metrics, we have traces, and we have logs all being shipped there. And that helps us really deep dive and understand when there's any type of performance regression, or incident happening in our cluster in real-time.

As far as what my critical tooling at Containous is, because I work in Marketing and because I work more in an educational-type atmosphere there, one of the tools that I have started to lean on heavily is something most people probably haven't heard of, and this is for managing the open-source community. It's something called Bitergia. And it's an analytics platform, but it helps me understand the health of the open-source community, and it helps me inform the engineering team of the activity around multiple projects, and who's contributing, and how long is it taking for issues and pull requests to be closed and merged? What's our ratio of pull requests and issues being closed for certain reasons. And these are all interesting business-y analytics that is important for our entire engineering organization to understand because we are an open-source company, and we rely heavily on our community for understanding the health of our business.

Emily: And speaking of, how can listeners connect with you?

Kevin: There's a couple different ways. One is through just plain old email. And that is kevin.crawley@containous—that’s C-O-N-T-A-I-N-O—dot U-S. And also through Twitter as well. And my handle is @notsureifkevin. It’s kind of like the Futurama, “Not sure if serious.” I mean, those are the two ways.

Emily: All right. Well, thank you so much. This was very, very interesting.

Kevin: Well, it was my pleasure. Thank you for taking the time to chat with me, and I look forward to listening to the podcast.

Emily: Thanks for listening. I hope you’ve learned just a little bit more about The Business of Cloud Native. If you’d like to connect with me or learn more about my positioning services, look me up on LinkedIn: I’m Emily Omier, that’s O-M-I-E-R, or visit my website which is emilyomier.com. Thank you, and until next time.

Announcer: This has been a HumblePod production. Stay humble.

View Details

The conversation covers:

  • An overview of Ravi’s role as an evangelist — an often misunderstood, but important technology enabler.
  • Balancing organizational versus individual needs when making decisions.
  • Some of the core motivations that are driving cloud native migrations today.
  • Why Ravi believes it in empowering engineers to make business decisions.
  • Some of the top misconceptions about cloud native. Ravi also provides his own definition of cloud native.
  • How cloud native architectures are forcing developers to “shift left.”

Links

  • https://harness.io/
  • Twitter: https://twitter.com/ravilach
  • Harness community: https://community.harness.io/
  • Harness Slack: https://harnesscommunity.slack.com/

Transcript

Emily: Hi everyone. I’m Emily Omier, your host, and my day job is helping companies position themselves in the cloud-native ecosystem so that their product’s value is obvious to end-users. I started this podcast because organizations embark on the cloud naive journey for business reasons, but in general, the industry doesn’t talk about them. Instead, we talk a lot about technical reasons. I’m hoping that with this podcast, we focus more on the business goals and business motivations that lead organizations to adopt cloud-native and Kubernetes. I hope you’ll join me.

Welcome to The Business of Cloud Native, I am your host Emily Omier. And today I'm chatting with Ravi Lachhman. Ravi, I want to always start out with, first of all, saying thank you—

Ravi: Sure, excited to be here.

Emily: —and second of all, I like to have you introduce yourself, in your own words. What do you do? Where do you work?

Ravi: Yes, sure. I'm an evangelist for Harness. So, what an evangelist does, I focus on the ecosystem, and I always like the joke, I marry people with software because when people think of evangelists, they think of a televangelist. Or at least that’s what I told my mother and she believes me still. I focus on the ecosystem Harness plays in. And so, Harness is a continuous delivery as a service company. So, what that means, all of the confidence-building steps that you need to get software into production, such as approvals, test orchestration, Harness, how to do that with lots of convention, and as a service.

Emily: So, when you start your day, walk me through what you're actually doing on a typical day?

Ravi: a typical day—dude, I wish there was a typical day because we wear so many hats as a start-up here, but kind of a typical day for me and a typical day for my team, I ended up reading a lot. I probably read about two hours a day, at least during the business day. Now, for some people that might not be a lot, but for me, that's a lot. So, I'll usually catch up with a lot of technology news and news in general. They kind of see how certain things are playing out.

So, a big fan of The New Stack big fan of InfoQ. I also like reading Hacker News for more emotional reading. The big orange angry site, I call Hacker News. And then really just interacting with the community and teams at large. So, I'm the person I used to make fun of, you know, quote-unquote, “thought leader.” I used to not understand what they do, then I became one that was like, “Oh, boy.” [laughs].

And so just providing guidance for some of our field teams, some of the marketing teams around the cloud-native ecosystem, what I'm seeing, what I'm hearing, my opinion on it. And that's pretty much it. And I get to do fun stuff like this, talking on podcasts, always excited to talk to folks and talk to the public. And then kind of just a mix of, say, making some sort of demos, or writing scaffolding code, just exploring new technologies. I'm pretty fortunate in my day to day activities.

Emily: And tell me a little bit more about marrying people with software. Are you the matchmaker? Are you the priest, what role?

Ravi: I can play all parts of the marrying lifecycle. Sometimes I'm the groom, sometimes I’m the priest. But I'm really helping folks make technical decisions. So, it’s go a joke because I get the opportunity to take a look at a wide swath of technology. And so just helping folks make technical decisions. Oh, is this new technology hot? Does this technology make sense? Does this project fatality? What do you think? I just play, kind of, masters of ceremony on folks who are making technology decisions.

Emily: What are some common decisions that you help people with, and common questions that they have?

Ravi: Lot of times it comes around common questions about technology. It's always finding rationale. Why are you leveraging a certain piece of technology? The ‘why’ question is always important. Let's say that you're a forward-thinking engineer or a forward-thinking technology leader.

They also read a lot, and so if they come across, let's say a new hot technology, or if they're on Twitter, seeing, yeah, this particular project’s getting a lot of retweets, or they go in GitHub and see oh, this project has little stars, or forks. What does that mean? So, part of my role when talking to people is actually to kind of help slow that roll down, saying, “Hey, what’s the business rationale behind you making a change? Why do you actually want to go about leveraging a certain, let's say, technology?”

I’m just taking more of a generic approach, saying, “Hey, what’s the shiny penny today might not be the shiny penny tomorrow.” And also just providing some sort of guidance like, “Hey, let's take a look at project vitality. Let's take a look at some other metrics that projects have, like defect close ratio—you know, how often it's updates happening, what's your security posture?” And so just walking through a more, I would say the non-fun tasks or non-functional tasks, and also looking about how to operationalize something like, “Hey, given you want to make sure you're maintaining innovation, and making sure that you're maintaining business controls, what are some best operational practices?” You know, want to go for gold, or don't boil the ocean, it’s helping people make decisive decisions.

Emily: What do you see as sort of the common threads that connect to the conversations that you have?

Ravi: Yeah, so I think a lot of the common threads are usually like people say, “Oh, we have to have it. We're going to fall behind if you don't use XYZ technology.” And when you really start getting to talking to them, it's like, let’s try to line up some sort of technical debt or business problem that you have, and how about are you going to solve these particular technical challenges? It's something that, of the space I play into, which is ironic, it's the double-edged sword, I call it ‘chasing conference tech.’ So, sometimes people see a really hot project, if my team implements this, I can go speak at a conference about a certain piece of technology.

And it's like, eh, is that a really rational reason? Maybe. It kind of goes into taking the conversation slightly somewhere else. One of the biggest challenges I think, let's say if you're kind of climbing the engineering ranks—and this is something that I had to do as I went from a junior to a staff to a principal engineer in my roles—with that it's always having some sort of portfolio. So, if you speak at a conference, you have a portfolio, people can Google your name, funny pictures of you are not the only things that come up, but some sort of technical knowledge, and sometimes that's what people are chasing. So, it's really trying to have to balance that emotional decision with what's best for the firm, what's best for you, and just what's best for the team.

Emily: That's actually a really interesting question is sometimes what's best for the individual engineer is not what's best for the organization. And when I say individual engineer, maybe it's not one individual, but five, or the team. How do you sort of help piece together and help people understand here's the business reason, that's organization-wide, but here's my personal motivation, and how do I reconcile these, and is there a way even to get both?

Ravi: There actually is a way to get both. I call it the 75/25 percent rule. And let's take all the experience away from the engineers, to start with a blank slate. It has to do with the organization. An organization needs to set up engineers to be successful in being innovative.

And so if we take the timeline or the scale all the way back to hiring, so when I like to hire folks, I always like to look at—my ratio is a little bit different than 75/25. I'm more of a 50/50. You bring 50 percent of the skills, and you'll learn 50 percent of the skills, versus more conservative organizations would say, “You know what? You have 75 percent of the skills, if you can learn 25 percent of the skills, this job would be interesting to you.” Versus if you have to learn 80 percent, it's going to be frustrating for the individual.

And so having that kind of leeway to make decisions, and also knowing that technical change can take a lot of time, I think, as an engineer, as an engineer—as talking software engineering professions as a whole, how do you build your value? So, your value is usually calculated in two parts. It’s calculated in your business domain experience and your technical skills. And so when you go project to project—and this is what might be more of, hey, if you’re facing too big of a climb, you'll usually change roles. Nobody is in their position for a decade. Gone are the days that you're a lifetime engineer on one project or one product.

It's kind of a given that you'll change around that because you're building your repertoire in two places: you're building domain experience, and you're building technical experience. And so knowing when to pick your battles, as cliche as that sounds, oh, you know what, this particular technology, this shiny penny came out. I seen a lot of it when Kubernetes came out, like, “Oh, we have to have it.” But—or even a lot of the cloud-native and container-based and all the ‘et cetera accessories’ as I call it, as those projects get steam surrounding it. It’s, “We have to have it.”

It's like, eh. It's good for resume building, but there's your things to do on your own also to learn it. I think we live in a day of open source. And so as an engineer, if I want to learn a new skill, I don't necessarily have to wait for my organization to implement it. I could go and play, something like Katacoda, I can go do things on my own, I can learn and then say, “You know what, this is a good fit. I can make a bigger play to help implement it in the organization than just me wanting to learn it.” Because a lot of the learning is free these days, which I think it's amazing. I know that was a long-winded answer. But I think you can kind of quench the thirst of knowledge with playing it on your own, and that if it makes sense, you can make a much better case to the business or to technology leadership to make change.

Emily: And what do you think the core business motivations are for most of the organizations that you end up talking to?

Ravi: Yeah, [unintelligible] core motivation to leveraging cloud-native technology, it really depends on organization to organization. I'm pretty fortunate that I get to span, I think, a wide swath of organization—so from startups to pretty established enterprises—I kind of talk about the pretty established enterprises. A lot of the business justification, it might not be a technical justification, but there's a pseudo technical business reason, a lot of times, though, I when I talk to folks, they're big concern is portability. And so, like, hey, if you take a look at the dollar and cents rationale behind certain things, the big play there is portability. So, if you're leveraging—we can get into the definition of what cloud-native resources are, but a big draw to that is being portable—and so, hopefully, you're not tied down to a single provider, or single purveyor, and you have the ability to move.

Now, that also ties into agility. Supposedly, if you're able to use ubiquitous hardware or semi-ubiquitous software, you were able to move a little bit faster. But again, what I usually see is folk’s main concern is portability. And then also with that is [unintelligible] up against scale. And so as—looking at ways of reducing resources, if you could use generics, you're able to shop around a little bit better, either internally or externally, and help provide scale for a softer or lesser cost.

Emily: And how frequently do you think the engineers that you talked to are aware of those core business motivations?

Ravi: Hmm, it really depends on—I’m always giving you the ‘depends’ answer because talking to a wide swath of folks—where I see there's more emotion involved in a good way if there's closer alignment to the business—which is something hard to do. I think it is slowly eroding and chipping away. I’ve definitely seen this during my career. It's the old stodgy business first technology argument, right. Like, modern teams, they're very well [unintelligible] together.

So, it's not a us versus them or cat versus dog argument, “Oh, why do these engineers want to take their sweet time?” versus, “Why does the business want us to act so fast?” So, having the engineers empowered to make decisions, and have them looked at instead of being a cost center, as the center of innovation is fairly key. And so having that type of rationale, like, hey, allowing the engineers to give input into feature development, even requirement development is something I've seen changed throughout my career. It used to be a very special thing to do requirements building, versus most of the projects that I've worked on now—as an engineer, we’re very, very well attuned to the requirements with the business.

Emily: Do you think there's anything that gets lost in translation?

Ravi: Oh, absolutely. As people, we're emotional. And so if we're all sum total of our experiences—so let's say if someone asked, Emily, you and I a question, we would probably have four different answers for that person, just because maybe we have differences in opinions, differences of sum totals of experience. And I might say, “Hey, try this or this,” and then you might say, “Try that or that.” So, it really depends.

Being lost in translation is always—it's been a fundamental problem in requirements gathering and it's continued to be a fundamental problem. I think just taking that question a step further, is how you go about combating that? I think having very shortened feedback cycles are very important. So, if you have to make any sort of adjustments, gone are the days I think when I started my career, waterfall was becoming unpopular, but the first project or two I was on was very waterfall-ish just because of the size of the project we worked on, we had to agree on lots of things; we were building something for six months. Versus, if you look at today, modern development methodologies like Agile, or Scaled Agile, a lot of the feedback happens pretty regularly, which can be exhausting, but decisions are made all the time.

Emily: Do you think in addition to mistranslations, do you think there are any misconceptions? And I'm talking about sort of on both sides of this equation, you know, business leaders or business motivations, and then also technologists, and let's refocus back to talk about cloud-native in particular. What sort of misconceptions do you think are sort of floating out there about cloud-native and what it means?

Ravi: Yeah, so what cloud-native means—it means something different to everybody. So, you listen to your podcasts for a couple episodes, if you asked any one of the guests the question, we all would give you a different answer. So, in my definition of cloud-native—and then I’ll get back to what some of the misconceptions are—I have a very basic definition: cloud-native means two pillars. It means your architecture, or your platform needs to be ephemeral, and it needs to be [indibited]. So, it needs to be able to be short-lived, and be consistent, which are two things that are at odds with each other.

But if you kind of talk to folks that, hey, they might be a little more slighted towards the business, they have this idea that cloud-native will solve all your problems. So, it reminds me a lot of big data back in the day. “Oh, if you have a Hadoop cluster, it will solve all of our logistics and shipping problems.” No. That's the technology. If you have Kubernetes, it will solve all of our problems. No. That's the technology. It's just a conduit of helping you make changes.

And so just making sure that understand that hey, cloud-native doesn't mean that you get the checkmark that, “Oh, you know what? We're stable. We're robust. We can scale by using all cloud-native technologies,” because cloud-native technologies are actually quite complicated. If you're introducing a lot of complexity to your architecture, does it make sense? Does that make sense? Does it give you the value you're looking for? Because at the end of the day, and this is kind of something, the older I get, the more I believe it, is that your customers don't care how you did something; they care what the result is. So, if your web application’s up, they don't care if you're running a simple LAMP stack, they just care that the application is up, versus using the latest Kubernetes stack, but using some sort of cloud-native NoSQL database, and we're using [Istio], and we’re using, pick your flavor du jour of cloud-native technology, your end customer actually doesn't care how you did it. They care what happened.

Emily: We can talk about misconceptions that other people have, but is there anything that continues to surprise you?

Ravi: Yeah, I think the biggest misconception is that there's very limited choice. And so I'll play devil's advocate, I think the CNCF, the Cloud Native Computing Foundation, there's lots of projects, I've seen the CNCF, they have something called the CNCF Landscape, and I seen it grow from 200 cards, it was 1200 cards at KubeCon, I guess, end of last year in San Diego, and it's hovering around 1500 cards. So, these cards means there's projects or vendors that play in this space. Having that much choice—this is usually surprising to people because they—if you're thinking of cloud-native, it's like saying Kleenex today, and you think of Kubernetes or other auxiliary product or project that surrounds that.

And a lot of misconception would be it's helping solve for complexity. It's the quintessential computer science argument. All you do in computer science is move complexity around like an abacus. We move it left to right. We’re just shifting it around, and so by leveraging certain technologies there's a lot of complication, a lot of burden that's brought in.

For example, if you want to leverage, let's say, a service Istio, Istio will not solve all your networking problems. In fact, it's going to introduce a whole set of problems. And I could talk about my biggest outage, and one of the things I see with cloud-native is a lot of skills are getting shifted left because you're codifying areas that were not codified before. But that's something I would love to talk about.

Emily: Tell me about your biggest outage that sounds interesting.

Ravi: Yeah, I didn't know how it would manifest itself. It’s ways, I think, until, like, years later that I didn’t have the aha moment. I used to think it was me, it probably still is me, but—so the year was 2013, and I was working for a client, and we were—it's actually a large news site—and so we were in the midst of modernizing their application, or their streaming application. And so I was one of the first applications to actually go to AWS. And so my background is in Java, so I have a Java software engineer or J2ED or JEE engineer, and having to start working more in infrastructure was kind of a new thing, so I was very fortunate up until 2013-ish up until this point that I didn't really touch the infrastructure. I was immune to that.

And now being more, kind of becoming a more senior engineer was in charge of the infrastructure for the application—which is kind of odd—but what ended up happening that—this is going to be kind of funny—since I was one of the first teams to go to AWS, the networking team wouldn't touch the configurations. So, when we were testing things, and [unintelligible] environments, we had our VPC CIDR rules—so the traffic rules—wide open. And then as we were going into production, there were rules that we had to limit traffic due to a CIDR so up until 2013, I thought a C-I-D-R like a CIDR was something you drink. I was like, “What? Like apple cider?” So, this shows you how much I know.

So, basically, I had to configure the VPC or Virtual Private Cloud networking rules. Finally, when we deployed the application, unknowing to myself, CIDR calculation is a significant digit calculation. So, the larger the number you divide by, the more IPs you let in. And so instead of dividing by 16, I divided by 8. I was like, “Oh, you’ll have a bigger number if you divide by a smaller number.”

I end up cutting off half the traffic of the internet when we deployed to production. So, that was a very not smooth way of doing something. But how did this manifest itself? So, the experts, who would have been the networking team, refused to look at my configuration because it was a public cloud. “Nope, you don’t have a slot in our data center, we look at it.” And poor me, as a JEE or J2EE engineer, I had very little experience networking.

Now, if you fast forward to what this means today, a lot of the cloud-native stack, are again, slicing and dicing these CNCF cards, a lot of this, you're exposing different, let's say verticals or dimensions to engineers that they haven't really seen before. A lot of its networking related a lot of it can be storage related. And so, as a software engineer, these are verticals that I’d never had to deal with before. Now, it's kind of ironic that in 2020, hey, yes, you will be dealing with certain configurations because, hey, it's code. So, it's shifting the burden left towards the developer that, “Oh, you know what, you know networking—” or, “You do need to know your app, so here's some Istio rules that you need to include in your packaging of your application.” Which folks might scratch your head.

So, yeah, again, it's like shifting complexity away from folks that have traditional expertise towards the developer. Now, times are changing. I seen a lot of this in years gone by, “Oh, no. These are pieces of code. We don't want to touch it.” Being more traditional or legacy operations team, versus today, everybody—it's kind of the merging of the two worlds. The going joke is all developers are becoming infrastructure engineers, and infrastructure engineers are becoming software engineers. So, it's the perfect blend of two worlds coming together.

Emily: That's interesting. And I now think I understand what you mean by skills shifting left. Developers have to know more, and more, and more. But I'm also curious, there's also people who talk about how Kubernetes, one of its failures is that it forces this shift left of skills and that the ideal world is that developers don't need to interact with it at all. That's just a platform team. What do you think about that?

Ravi: These are awesome questions. These are things I'm very passionate about. I definitely seen the evolution. So, I've been pretty fortunate that I was jumping on the application infrastructure shift around 2014, 2015, so right when Kubernetes was coming of age. So, most of my background was in distributed systems.

So, I'm making very large distributed Java applications. And so when Kubernetes came out, the teams that I worked on, the applications that were deployed to Kubernetes were actually owned by the app dev team. The infrastructure team wouldn't even touch the Kubernetes cluster. It was like, “Oh, this is a development tool. This is not a platform tool.”

The platform teams that I were interacting with 2015, 2016, as Kubernetes became more popular than ever, they were the legacy—well, hate to say legacy because it’s kind of my background too—they were the remaining middleware engineers. We maintained a web server cluster, we maintained the message broker cluster, we maintained XYZ distributed Java infrastructure cluster. And so when looking at a tool like Kubernetes, or even there were different platforming services, so the paths I've leveraged early, or mid-2010s was Red Hat OpenShift, before and after the Kubernetes migration inside of OpenShift. And so looking at a different—how teams are set up, it used to be, “Oh, this is an app dev item. This is what houses your application.”

Versus today, because the workloads are so critical that are going on to say platforms such as Kubernetes, it was that you really need that system engineering bubble of expertise. You really need those platform engineers to understand how to best scale, how to best purvey, and maintain a platform like Kubernetes. Also, one of the odd things are—going back to your point, Emily, like, hey, why things were tossed over either to the development team or going back to a developing software engineer myself, do we care what the end system is?

So, it used to be, I'll talk about Java-land here for a minute, give you kind of long-winded answer of back in Java land, we really used to care about the target system, not necessarily for an application that have one node, but if we had to develop a clustered application. So, we have more than one node talking to each other, or a stateful application, we really had start developing to a specific target system. Okay, I know how JBoss WildFly clusters or I know how IBM WebSphere or WebLogic clusters. And so when we're designing our applications, we had to make sure that we play well into those clustering mechanisms. With Kubernetes, since it's generic, you don't necessarily have to play into those clustering mechanisms because there's a basic understanding. But that's been the biggest Achilles heel in Kubernetes. It wasn't designed for those type of workloads, stateful workloads that don't like dying very often. That's kind of been the push or pull. It's just a tool, there's a lot of generic, so you can assume that the target platform will handle a certain way. And you're slowly start backing off the case that you're building to a specific target platform. But as Kubernetes has evolved, especially with the operator framework, you actually are starting to build to Kubernetes in 2018, 2019, 2020.

Emily: It actually brought up a question for me that, at risk of sounding naive myself, I feel like I never meet anybody who introduces themselves as a platform engineer. I meet all these developers, everyone's a developer evangelist, for example, or their background is as a developer, I feel like maybe once or twice, someone has introduced themselves as, “I’m a platform engineer,” or, “I’m an operations specialist.” I mean, is that just me? Is that a real thing?

Ravi: They’re very real jobs. I think… it's like saying DevOps engineer, it means something else to who you talk to you. So, I'll harp on, like ‘platform engineer.’ so kind of like, the evolution of the platform engineer, if you would have talked to me in 2013, 2014, “Hey, I'm a platform engineer,” I would think that you're a software engineer focused on platform tools. Like, “Hey, I focus on authentication, authorization.”

You're building—let's say we had a dozen people on this call and we're working for Acme Incorporated, there's modules that transcend every one of our teams. Let's say logging, or let’s say login, or let's say, some sort of look and feel. So, the platform engineer or the platform engineering development focused platform engineering team would make common reusable modules throughout. Now, with the great rise of platforms as a service, like PCF, and OpenShift, and DCOS, they became kind of like a shift. The middleware engineers that were maintaining the message broker clusters, maintaining your web application server clusters, they’re kind of shifting towards one of those platforms.

Even today, Kubernetes, pick your provider du jour of Kubernetes. And so those are where the platform engineers are today. “Hey, I'm a platform engineer. I focus on OpenShift and Kubernetes.” Usually, they're very vertically focused on one or more specific platforms. And operations folks can ride very big gamut. Usually, if you put, “operations” in quotes, usually they’re systems or infrastructure engineers that are very focused on the infrastructure where the platform’s run.

Emily: I'm obviously a words person, and it just seems like there's this vocabulary issue where everybody knows what a developer is, and so it's easy to say, “Oh, I'm a developer.” But then everything else that's related to engineering, there's not quite as much specificity, precisely because you said everybody has a slightly different understanding. It's kind of interesting.

Ravi: Yeah, it's like, I think as a engineer, we're not one for titles. So, I think a engineer is a engineer. I think if you asked most engineers, it’s like, “Yeah, I’m a engineer.” It's so funny, a good example of that is Tim Berners-Lee, the person who created WWW, the World Wide Web. If you looked at his LinkedIn, he just says he's a web developer. And he invented WWW. So, usually engineering-level folks, you're not—at least for myself—is not one for title.

Emily: The example that you gave regarding the biggest outage of your career was basically a skills problem. Do you think that there's still a skills or knowledge issue in the cloud-native world?

Ravi: Oh, absolutely. We work for incentivization. You know, my mortgage is with PNC, and they require a payment every month, unfortunately. So, I do work for an employer. Incentivization is key. So, kind of resume chasing, conference chasing there's been some of that in the cloud-native world, but what ends up happening more often than not is that we're continuously shifting left.

A talk I like to give is called, “The Engineering Burden is on the Rise.” And taking a look at what, let's say, a software engineer was required to do in 2010 versus what a software engineer is required to do today in 2020. And there's a lot more burden in infrastructure that, as a software engineer you didn't have to deal with. Now, this has to do with two things, or actually one particular movement. There's a movie company, or a video company in Los Gatos, California, and there's a book company in South Lake Union in Seattle.

And so these two particular companies given the rise of what's called a full lifecycle developer. Basically, if you run it, or if you operate—you operate what you run, or if you write it, you run it. So, that means that if you write a piece of code, you're in charge of the operations. You have support, you're in charge of the SLAs, SLOs, SLIs. You're ultimately responsible if a customer has a problem.

And can you imagine the number of people, the amount of skill set that requires? There's this concept of a T-shaped skill that you have to have experience in so many different platforms, that it becomes a very big burden. As an engineer, I don't envy anybody entering a team that's leveraging a lot of cloud-native technology because most likely a lot of that onus will fall on the software engineer to create the deployable, to create how you build it, to fly [unintelligible] in your CI stack, write the configuration that builds it, write the configuration deploys it, write the networking rules, write how you test it, write the login interceptors. So, there's a lot going on.

Emily: Is there anything else that you want to add about your experience with cloud-native that I haven't really thought to ask, yet?

Ravi: It's not all doom and gloom. I'm very positive on cloud-native technologies. I think it's a great equalizer. You're kind of going back—this might be a more intrinsic, like a 30-second answer here. If you taking back that I wanted to learn certain skills in 2010, I basically had to be working for a firm. So, 2010, I was working for IBM. So, there's certain distributed Java problems I wanted to solve. I basically had to be working for a firm because the software licensing costs were so expensive, and that technology wasn't very democratized.

Looking at cloud-native technology today, there's a big, big push for open source, which open source is R&D methodology. That's what open source is, it helps alleviate some sort of acquisition—but not necessarily adoption—problems. And you can learn a lot. Hey, you could pick up any project and just try to learn, try to run it. Pick up these particular distributed system skills that were very guarded, I would say, a decade ago, it's being opened up to the masses. And so there's a lot to drink from, but you can drink as much as you want from the CNCF or the cloud-native garden hose.

Emily: Do you have a software engineering tool that you cannot live without?

Ravi: Recently, because I deal in a lot of YAML, I need a YAML linter. So, YAML is a space-separated language. As a human, I can't tell you what spaces are. Like, you know, if you have three spaces, and the next line you have four spaces. So, I use a YAML linter. It puts periods for me, so I can count them because it's been multiple times that my demo is not syntactically correct because I missed a space and I can't see it on my screen.

Emily: And how can listeners connect with you?

Ravi: Oh, yeah. You can hit me up on Twitter @ravilach, R-A-V-I-L-A-C-H. Or come visit us at Harness at www.harness.io. I run the Harness community, so community.harness.io. We have a Slack channel and a Discourse, and always excited to interact with people.

Emily: Thanks for listening. I hope you’ve learned just a little bit more about the business of cloud-native. If you’d like to connect with me or learn more about my positioning services, look me up on LinkedIn: I’m Emily Omier, that’s O-M-I-E-R, or visit my website which is emilyomier.com. Thank you, and until next time.

Announcer: This has been a HumblePod production. Stay humble.

View Details

The conversation covers:

  • Some of the pain points and driving factors that led Jón and his partners to launch Garden. Jon also talks about his early engineering experiences prior to Garden.
  • How the developer experience can impact the overall productivity of a company, and why companies should try and optimize it.
  • Kubernetes shortcomings, and the challenges that developers often face when working with it. Jón also talks about the Kubernetes skills gap, and how Garden helps to close that gap.
  • Business stakeholder perception regarding Kuberentes challenges.
  • The challenge of deploying a single service on Kubernetes in a secure manner — and why Jón was surprised by this process.
  • How the Kubernetes ecosystem has grown, and the benefits of working with a large community of people who are committed to improving it.
  • Jón’s multi-faceted role as CEO of Garden, and what his day typically entails as a developer, producer, and liaison.
  • Garden’s main mission, which involves streamlining end-to-end application testing.

Links:

  • Company site: https://garden.io/
  • Twitter: https://twitter.com/jonedvald
  • Kubernetes Slack: https://slack.k8s.io/

Transcript:

Emily: Hi everyone. I’m Emily Omier, your host, and my day job is helping companies position themselves in the cloud-native ecosystem so that their product’s value is obvious to end-users. I started this podcast because organizations embark on the cloud naive journey for business reasons, but in general, the industry doesn’t talk about them. Instead, we talk a lot about technical reasons. I’m hoping that with this podcast, we focus more on the business goals and business motivations that lead organizations to adopt cloud-native and Kubernetes. I hope you’ll join me.

Emily: Welcome to The Business of Cloud Native. I'm your host Emily Omier. And today I'm chatting with Jón Eðvald. And, Jón, thank you so much for joining me.

Jón: Thank you so much for having me. You got the name pretty spot on. Kudos.

Emily: Woohoo, I try. So, if you could actually just start by introducing yourself and where you work in Garden, that would be great.

Jón: Sure. So, yeah, my name is Jón, one of the founders, and I’m the CEO of Garden. I've been doing software engineering for more years than I'd like to count, but Garden is my second startup. Previous company was some years ago; dropped out of Uni to start what became a natural language processing company. So, different sort of thing than what I'm doing now.

But it's actually interesting just to scan through the history of how we used to do things compared to today. We ran servers out of basically a cupboard with a fan in it, back in the day, and now, things are done somewhat differently. So, yeah, I moved to Berlin, it's about four years ago now, met my current co-founders. We all shared a passion and, I guess to some degree, frustrations about the general developer experience around, I guess, distributed systems in general. And now it's become a lot about Kubernetes these days in the cloud-native world, but we are interested in addressing common developer headaches regarding all things microservices.

Testing, in particular, has become a big part of our focus. Garden itself is an open-source product that aims to ease the developer experience around Kubernetes, again, with an emphasis on testing. When we started it, there wasn't a lot of these types of tools around, or they were pretty early on. Now there's a whole bunch of them, so we're trying to fit into this broad ecosystem. Happy to expand on that journey. But yeah, that's roughly—that's what Garden is, and that’s… yeah, a few hop-skips of my history as well.

Emily: So, tell me a little bit more about the frustration that led you to start Garden. What were you doing, and what were you having trouble doing, basically?

Jón: So, when I first moved to Berlin, it was to work for a company called Clue. They make a popular period tracking app. So, initially, I was meant to focus on the data science and data engineering side of things, but it became apparent that there was a lot of need for people on the engineering side as well. So, I gravitated into that and ended up managing the engineering team there. And it was a small operation. We had more than a million daily active users yet just a single back end developer, so it was bursting at the seams.

And at the time running a simple Node.js backend on Heroku, single Postgres database, pretty simple. And I took that through—first, we adopted containers and moved into Docker Cloud. Then Docker Cloud disappeared, or was terminated without—we had to discover that by ourselves. And then Kubernetes was manifesting as the de facto way to do these things. So, we went through that transition, and I was kind of surprised. It was easy enough to get going and get to a functional level with Kubernetes and get everything running and working. The frustration came more from just the general developer experience and developer productivity side. Specifically, we found it very difficult to test the whole application because we had, by the end of that journey, a few different services doing different things. And for just the time you make a simple change to your code to it actually having been built, deployed, and ultimately tested was a rather tedious experience. And I found myself building tools, bespoke tools to be able to deal with that, and that ended up being sort of a janky prototype of what Garden is today. And I realized that my passion was getting the better of me, and we wanted to start a company to try and do better.

Emily: Why do you think developer experience matters?

Jón: Beyond just the, kind of, psychological effect of having to have these long and tedious feedback loops—just as a developer myself, it kind of grinds and reduces the overall joy of working on something. But in more concrete material terms, it really limits your productivity. You basically, you take—if your feedback loop is 10 times longer than it should be, that exponentially reduces the overall output of you as an individual or your team. So, it has a pretty significant impact on just the overall productivity of a company.

Emily: And, in fact, it seems like a lot of companies move to Kubernetes or adopt distributed systems, cloud-native in general, precisely to get the speed.

Jón: And, yeah, that makes sense. I think it's easy to underestimate all the, what are often called these day-two problems, when—so, it's easy enough to grok how you might adopt Kubernetes. You might get the application working, and you even get to production fairly quickly, and then you find that you've left a lot of problems unsolved, that Kubernetes by itself doesn't really address for you. And it's often conflated by the fact that you may be actually adopting multiple things at the same time. You may be not only transitioning to Kubernetes from something analogous, you may be going from simpler, bespoke processes, or you might have just a monolith that didn't really have any complicated requirements when it comes to dev tooling and dev setups. So, yeah, you might be adopting microservices, containers, and Kubernetes all at the same time, and I can say a lot of good things about Kubernetes, but it leaves, definitely, a lot of gaps in terms of the experience, especially for the application developer.

So, you have these different kinds of developers, and we're kind of gradually, and somewhat clumsy sometimes, moving from a world where would have the application developer, and then you would have the operators, and you would kind of throw things over a fence. And now we're trying to fuse these a little bit, and you end up with this combination of different people with different strengths, and different skill sets and the [00:07:03 infra], and often termed DevOps engineers—I know that makes some people cringe. It's just what we see out there, that's a job title that we’re seeing more and more—they are at home with this sort of a thing, and they maybe have time and patience to tinker with something like Kubernetes, whereas your application developer, they just want to get that feature out. So, you end up with these two different disciplines where one is basically trying to unblock the other. And we’re definitely made progress. And I like to think what we are building helps, and all the other tools that are kind of overlapping in our space. But I still see a lot of this frustration with both having to adopt this more DevOps mindset, and also just the frustrations with Kubernetes specifically, and all the actual technologies that are in play these days.

Emily: And to what extent do you think this sub-optimal developer experience is about a skills gap?

Jón: I think that's definitely a part of it. I mean, we have both current and prospective customers in all kinds of different companies coming from different backgrounds, and some more old school and rigid than others. And I would say, for companies where the divide was really crisp—like, you had the application developer, and then you had the operator—then basically the mentality, understandably, for the application developer, they just want to work on their code, and not have to worry about networking ports, and ingresses, and all the low-level primitives of Kubernetes. It’s a whole new thing to learn, and the perception is that it's more getting in the way of them getting stuff done. They don't really need the power for their immediate needs. The power is maybe something that the operator really enjoys, and has a strong need for.

Emily: Do you think tools like Garden help close the skills gap?

Jón: I think so. It definitely makes it more palatable at the start; it's easier to get going, you get some guide rails, and you get some higher-level abstractions that are maybe familiar to you, so you're not completely overwhelmed by all the different options, and flags and whatnot that you need for all the Kubernetes, YAML, et cetera, and you can ease into it. So, I think that's really helpful. Another thing that I think is really helpful is, so instead of this separation—[00:09:33 crowbar] separation between developer and operator, people coming from the operations side, part of their role now is to empower the developer, so the developer becomes their customer. I think it, in some sense, always was that way, but now they are more focused on the developer experience, and unblocking, and creating, often, these internal tools and internal abstractions that buffer this skills gap, I guess, a little bit.

And we see all kinds of variations on how people actually go about it. With Garden, what we hope to achieve is for—there's usually a team that is responsible for the developer experience of an organization, and they can just hand tools like that to the team. And maybe there's a little bit of a dotted line between what the application developer is responsible for and what they should know, but they should still be familiar enough with the concept so that they can start to perceive the infrastructure and the general systems architecture of the application as part of the application and not this separate concern. Yeah, I mean, I definitely would hope that were helpful in this regard.

Emily: Yeah, and I was going to ask, actually, what your opinion about this is, you know about how familiar an application developer actually needs to be with Kubernetes.

Jón: I think unless there is substantial investment at their company in abstracting these things away from the developer, they need to have at least a cursory understanding of how all this set up. Maybe they're not responsible for actually making the Kubernetes manifest, maybe they're working in some kind of a templated fashion, or they have a little bit of a buffer between the low-level configuration of everything and what they're working on in a day-to-day, but they'll get stuck pretty quickly and it will create a lot of friction if they just don't know at all how the thing works. I think it's also just important for any, kind of, distributed systems developers to understand what the primitives are, even if you don't know which flag to use for this and that to actually achieve what you want to achieve, understanding the dynamics and understanding—you know, as a developer, you're meant to understand basically how a computer works, you know, the general layout of a computer, CPU or memory, and you have some notion of how the network works. This is maybe a layer on top of that similar to how you are familiar with how an operating system works, you should have basic familiarity over how Kubernetes works and just the general layout of a cloud-native ecosystem. Which I think is reasonable.

Emily: And then when we think about business stakeholders in a company, so people who are outside of the IT department, how do they perceive this challenge? Do they understand how developers might struggle with moving to Kubernetes? Do they understand the stakes of getting developer experience right?

Jón: I think a lot of the time they find out in the process. I think, for better or for worse, this cloud-native ecosystem was kind of jumpstarted, and people were sold on the idea a little bit before it was actually palatable for day-to-day development and actually running in production. So, you have an ecosystem running around trying to make this into a pleasant experience, but part of that meant that there have been—especially the early adopters, I can only imagine, but even companies adopting Kubernetes now, yeah, they maybe have the wrong picture of what it actually takes to do this successfully, and we've definitely seen—I mean, I'm kind of in that group myself. I adopted Kubernetes fairly early on. Well,-ish; it was about three years ago.

And once you get past day one, as it were, there is a lot of surprises, and I think that could have done with a little bit more education, and maybe the motivation to get people on board kind of clouded that a little bit. No pun intended. So, yeah, I think it's more common than not that people have had to find that out along the way. Maybe that's probably getting better now. I think people have a—both the tooling is actually better so that gap isn't as wide as it was, but also to general awareness, I think there's enough stories out there where people have really struggled with the adoption, and struggled with productivity somewhere along the way, or after the adoption process. Yeah. Does that answer the question? [laughs].

Emily: Absolutely. And actually, I'm curious, what were some of the surprises for you?

Jón: I'm still surprised every now and then. I've been working, I’ve been developing against Kubernetes, full-time pretty much—well, to the extent that I can develop full time over the last couple of years—and worked with it as an end-user before that, and I'm still finding new and interesting ways in which something can fail. And just a simple example, just the act of deploying a single service onto Kubernetes, you need to write up a few different YAML files. Doing that in a secure manner is not exactly obvious, and is nowhere near secure by default, so you have to, kind of, discover that, or pay some security providers to really help you with that. The number of ways in which a service deployment can fail is just—it's still—it keeps rising. [laughs]. I found a new one just yesterday. Like, “Oh, actually, here's an interesting condition. I hadn't come across that before.”

So, the abstractions are fairly low-level, and you have to be able to navigate that. And I would much prefer higher-level abstractions that encapsulate all the—even if it's just encapsulating all the different failure modes of starting a container, that would be super helpful. But I think that's one thing that people keep stumbling on. And just to deploy a single service, there's a lot of different bells and whistles that you need to be kind of faintly aware of, and they—you will keep discovering them as you keep working with it. It's been both an interesting experience, and it's definitely put a lot of life into this developer ecosystem, but it's also been quite frustrating since myself and a lot of other developers out there are stuck having to deal with this and discover all these different things gradually. Does that make sense?

Emily: Absolutely, but I'm going to ask another question, which is, what about pleasant surprises?

Jón: I think I mentioned one earlier. One of the pleasant surprises is how quickly this ecosystem has formed, and a huge community of people, and really putting a lot of effort into making the whole experience better, and adding any functionality that doesn't come out of the box, as it were. I think it's been really interesting to observe that ecosystem manifest. Yeah, I think that's probably the most interesting positive thing that comes to mind. I think it's been interesting to see all kinds of different, ostensibly competing companies work together in many ways.

Just the general power of the open-source movement has definitely come to light which has been great. I feel like the attitudes just in our user community, and the user community generally—just go to the Kubernetes Slack, or you have discussions on GitHub issues or all kinds of, you know, one or more of these open-source projects and ecosystem. Like, just general morale, and helpfulness, and collaborative nature has been really interesting to watch, and I’ve really enjoyed that.

Emily: I'm going to ask, just sort of a very different question, which is, I always like to ask my guests what their job actually entails. Everyone has a job title, and it doesn't always provide lots of visibility into what you actually do during the day. So, what does somebody like you, who's the CEO of a startup, what do you actually do when you get into the office? What does your day actually look like?

Jón: So, it has evolved a lot. Let's put it that way. And I wear many hats for sure. So, my background is I'm an engineer, that's kind of my comfort zone. I can write code and I can think about systems and all of that jazz, and that comes naturally to me. So, in my CEO role, I've been learning a lot on the go.

So, these days, a lot of my time, I've spent maybe 30, 40 percent of my time coding these days, which is just about average. Sometimes it's a little bit more, sometimes it's almost not at all. And that's probably going to fluctuate and, over time, decrease. I think the most important part of my role is to generally engage with the ecosystem, and that involves reading, talking to people, talking to investors, talking to other people in my industry, talking to customers, kind of trying to build a mental map of this world we live in, and try and figure out how that translates into our product roadmap for us, and try and communicate what's important, what's urgent for the team to think about. So, I'm kind of this liaison between the external world and the internal world that is our team in many different ways. I think that consumes a lot of my time.

And then working on content, working with our marketing manager, joining sales calls, supporting customers, working with our designer trying to figure out what's our brand identity, and how should that evolve. Yeah, I’m kind of a nexus of a lot of different things, which I really enjoy. Sometimes exhausting, don't get me wrong. But yeah, I like to think that I'm connecting dots.

Emily: So, I wanted to also ask a little bit more about what Garden does, and particularly how is the testing that you do, how is it different from what you would get in just a normal CI tool?

Jón: Right. So, testing is definitely—it was always part of the picture when we started. I think when we started working on Garden, we maybe tried to solve too many problems. That's pretty common startup mistake, I guess. And we're focusing more and more on testing. And so, I'm happy to elaborate on that.

So, the way people do testing today—and this I derive from both personal experience and talking to a lot of our current and prospective customers—it's very common to, well A) not to it in any meaningful way. That's surprisingly common still, or to unit test the individual components of an application. But that falls short in a number of different ways when working with just microservices in general. And in these cloud-native scenarios, you tend to have a lot of different moving parts. And what we're trying to address is how easily you can test your application end-to-end; how you can spin up your own ad hoc instance of an application, the full application with all the different services, even the ones you do not working on; and how you can write tests such that they actually test the interaction between services and not through mocking and stubbing them, or bending over backwards to do these contract testing and things like that.

That certainly helps mitigate these things, but the cloud-native ecosystem actually bring—and Kubernetes, and just the declarative nature of how we define and deploy applications today—actually creates a lot of opportunities. You can now feasibly spin up a whole instance of an application with all different services that comprise your back end, and you can run tests from within there. And Garden really helps to make that efficient and easy to use for developers. So, the comparison to CI naturally comes up, and Garden in a certain sense blurs the line between a developer tool that you use in the inner loop or locally from your laptop, and the process that happens in your CI. We don't aim to replace CI; that's a pretty saturated market. We just want to focus on the testing and previewing side of it.

And what Garden does, which is unique, is that it allows you to describe every part of your stack, how it's built, how is deployed, and how it's tested, and we compose this into what we call the Stack Graph. So, we have this directed graph of all the different things that need to happen between just a bunch of Git repositories and a fully running, tested system. And where that comes in really handy is when you make a small change—you push a pull request or you're just testing locally by yourself—we know which parts of your stack actually are affected by that change, so you don't need to run a full battery of all of your different tests, you end-to-end tests. You can actually just automatically trigger the few parts that need to be rebuilt, redeployed, and retested. So, it's making it viable to actually run integration tests and end-to-end tests as part of your development workflow, and it also makes it much easier for the developer to cross between the inner loop—working on the code themselves—and getting it to pass through CI because that's often an arduous process, just getting things to pass testing in CI, because it may work fine on your laptop, or you may only have a small subset of the system that you're working on, and then you have to make another commit and push to get, and then the CI pipeline runs again. That can take an arbitrarily long amount of time. And instead, you can just run Garden test, or run a specific test suite from your laptop and have it work much the same way in CI as when you're developing locally.

Emily: Do you have anything that you can think of that you'd like to add about this general topic of the surprises and misconceptions that come up around cloud-native and Kubernetes?

Jón: It's generally been a pleasant experience. And it's been really motivating to see all this activity in this space. There's definitely lots of activity, lots of small startups and big companies trying to figure out better ways to do things. And I think it's fun to be a part of it. I think, who's to say how it's all going to shake out, but I don't remember it being quite like this pre-cloud-native, and I'm very much looking forward to, at the end of this journey, that we as an ecosystem kind of are going through that we will have something that really materially changes how we develop distributed systems, not just in the context of Kubernetes—because that's a specific technology that works for a specific type of customer—but rather taking it further and keeping this collaborative note, even when we go past what's the technology of the day.

Emily: All right, so one bonus question, what is your favorite engineering tool that you could not live without? Or I should say that you could not do your work without? Maybe you could live without it.

Jón: Well, yeah, exactly. Yeah. I think probably the most central tool that I have in terms of my development is Google. The thing is, as a developer, there's just no way you can know all the things. You don't remember every API. Like, I keep looking at the same thing over and over again, and being able to have all this information just at my fingertips, I think that's something that I would—well actually both have a hard time living without, and engineering without.

And notable mentions, I've really enjoyed working with [00:25:33 VS Code] as of late. We code in TypeScript, which is slightly unusual, but I've also become a fan of that. There's probably a number of smaller tools that I'm forgetting to mention that I use all the time, but just being able to find whatever I need at any time, that's probably the most revolutionary thing to come up over the last decade.

Emily: Where could listeners connect with you or follow you?

Jón: So, I'm on Twitter, not the most active on Twitter, but I'm always—if you follow me, I probably just follow you back and try and get into the conversation that way. You can reach me there directly. Also, if you just want to ping me—so on Twitter, @jonedvald, just my name as it's printed. You can find the community for Garden itself through our website, Garden.io. We have a Slack channel on the Kubernetes Slack, just #Garden. And yeah, please reach out. I'd love to chat. If you're interested in using Garden or not, I'm always trying to be more engaged with the community, and that doesn't have to be Garden-related necessarily.

Emily: Well, thank you so much for joining us on The Business of Cloud Native.

Jón: Thank you so much for having me.

Emily: Thanks for listening. I hope you’ve learned just a little bit more about the business of cloud native. If you’d like to connect with me or learn more about my positioning services, look me up on LinkedIn: I’m Emily Omier, that’s O-M-I-E-R, or visit my website which is emilyomier.com. Thank you, and until next time.

Announcer: This has been a HumblePod production. Stay humble.

View Details

Some of the highlights of the show include:

  • The challenges that CERN was facing when storing, processing, and analyzing data, and why it pushed them to think about containerization.
  • CERN’s evolution from using mainframes, to physical commodity hardware, to virtualization and private clouds, and eventually to containers. Ricardo also explains how the migration to containerization and Kubernetes was started.
  • Why there was a big push from groups that focus on reproducibility to explore containerization.
  • How end users have responded to Kubernetes and containers. Ricardo talks about the steep Kubernetes learning curve, and how they dealt with frustration and resistance.
  • Some of top benefits of migrating to Kubernetes, and the impact that the move has had on their end users.
  • Current challenges that CERN is working through, regarding hybrid infrastructure and rising data loads. Ricardo also talks about how CERN optimizes system resources for their scientists, and what it’s like operating as a public sector organization.
  • How CERN handles large data transfers.

Links:

  • Email:ricardo.rocha@cern.ch
  • Twitter: https://twitter.com/ahcorporto
  • CERN

Transcript

Emily: Hi everyone. I’m Emily Omier, your host, and my day job is helping companies position themselves in the cloud-native ecosystem so that their product’s value is obvious to end-users. I started this podcast because organizations embark on the cloud naive journey for business reasons, but in general, the industry doesn’t talk about them. Instead, we talk a lot about technical reasons. I’m hoping that with this podcast, we focus more on the business goals and business motivations that lead organizations to adopt cloud-native and Kubernetes. I hope you’ll join me.

Emily: Welcome to the Business of Cloud Native. I'm your host, Emily Omier, and today I'm here with Ricardo Rocha. Ricardo, thank you so much for joining us.

Ricardo: It's a pleasure.

Emily: Ricardo, can you actually go ahead and introduce yourself: where you work, and what you do?

Ricardo: Yeah, yes, sure. I work at CERN, the European Organization for Nuclear Research. I'm a software engineer and I work in the CERN IT department. I've done quite a few different things in the past in the organization, including software development in the areas of storage and monitoring, and also distributed computing. But right now, I'm part of the CERN Cloud Team, and we manage the CERN private cloud and all the resources we have. And I focus mostly on networking and containerization, so Kubernetes and all these new technologies.

Emily: And on a day to day basis, what do you usually do? What sort of activities are you actually doing?

Ricardo: Yeah. So, it's mostly making sure we provide the infrastructure that our physics users and experiments require, and also the people on campus. So, CERN is a pretty large organization. We have around 10,000 people on-site, and many more around the world that depend on our resources. So, we operate private clouds, we basically do DevOps-style work. And we have a team dedicated for the Cloud, but also for other areas of the data center. And it's mostly making sure everything operates correctly; try to automate more and more, so we do some improvements gradually; and then giving support to our users.

Emily: Just so everyone knows, can you tell a little bit more about what kind of work is done at CERN? What kind of experiments people are running?

Ricardo: Our main goal is fundamental research. So, we try to answer some questions about the universe. So, what's dark matter? What's dark energy? Why don't we see antimatter? And similar questions. And for that, we build very large experiments.

So, the biggest experiment we have, which is actually the biggest scientific experiment ever built, is the Large Hadron Collider, and this is a particle accelerator that accelerates two beams of protons in opposite directions, and we make them collide at very specific points where we build this very large physics experiments that try to understand what happens in these collisions and try to look for new physics. And in reality, what happens with these collisions is that we generate large amounts of data that need to be stored, and processed, and analyzed, so the IT infrastructure that we support, it’s larger fraction dedicated to this physics analysis.

Emily: Tell me a little bit more about some of the challenges related to processing and storing the huge amount of data that you have. And also, how this has evolved, and how it pushed you to think about containerization.

Ricardo: The big challenge we have is the amount of data that we have to support. So, these experiments, each of the experiments, at the moment of the collisions, it can generate data in the order of one petabyte a second. This is, of course, not something we can handle, so the first thing we do, we use these hardware triggers to filter this data quite significantly, but we still generate, per experiment, something like a few gigabytes a second, so up to 10 gigabytes a second. And this we have to store, and then we have large farms that will handle the processing and the reconstruction of all of this. So, we've had these sort of experiments since quite a while, and to analyze all of this, we need a large amount of resources, and with time.

If you come and visit CERN, you can see a bit of the history of computing, kind of evolving with what we used to have in the past in our data center. But it's mostly—we used to have large mainframes, that now it's more in the movies that we see them, but we used to have quite a few of those. And then we transitioned to physical commodity hardware with Linux servers. Eventually introduced virtualization and private clouds to improve the efficiency and the provisioning of these resources to our users, and then eventually, we moved to containers and the main motivation is always to try to be as efficient as possible, and to speed up this process of provisioning resources, and be more flexible in the way we assign compute and also storage.

What we've seen is that in the move from physical to virtualization, we saw that the provisioning and maintenance got significantly improved. What we see with containerization is the extra speed in also deployment and update of the applications that run on those resources. And we also see an improving resource utilization. We already had the possibility to improve quite a bit with virtualization by doing things like overcommit, but with containers, we can go one step further by doing more efficient resource sharing for the different applications we have to run.

Emily: Is the amount of data that you're processing stable? Is it steadily increasing, have spikes, a combination?

Ricardo: So, the way it works is, we have what we call ‘beam’ which is when we actually have protons circulating in the accelerator. And during these periods, we try to get as much collisions as possible, which means as much data as possible because this is the main ingredient we need to do physics analysis. And then with the current experiment, actually, we are quite efficient. So, we have a steady stream of data we generate from raw data is something like 70 petabytes a year, and then out of that data, we generate a lot more data be it like pre-processing, or physics analysis. So, in terms of data, we can pretty much calculate how much storage we need. What we see spikes is in the computing part.

Emily: And is the amount of data, are you constantly having to increase your data storage capacity?

Ricardo: Yeah, so, yeah, that's one aspect is we never—from the raw data, we keep it forever. So, actually, we need to increase what we do, which is kind of particular to the way CERN has been doing things is we rely a lot on tape for archival. And this has allowed us to—because states still evolve, and the capacity of the tape still evolves, this actually allows us to increase the capacity on our data center without having to get more space or more robots. We just increase the capacity of the tape storage. And then for the on-disk storage, though, we do have to increase capacity regularly.

Emily: Who spearheaded the push to containerization and Kubernetes? Was it within the IT? Who in the organization was really pushing for it?

Ricardo: Yeah. That's an interesting question because it happens gradually in different places. I wouldn't say it's a decision taken from top; bottom. It's more like a movement that starts slowly. And there's two main triggers, I would say.

One of them is from the IT department, the people run the services because some of our engineers are exposed to these new technologies frequently, and there's always someone that sees some benefit, does a prototype, and then presents to their colleagues and this kind of trigger. So, it's more on the sense of simplifying operations, deployment, all the monitoring that is required, and also improve resource usage. So, this was one of the triggers.

The other trigger is actually from our end users. One big aspect of science, and specifically here—physics—is that there's a big need to share a setup. When someone performs some analysis, it's quite common that they want to share the exact same setup with their colleagues, and make sure that this analysis is reproducible. So, there was a huge push from groups that focus on reproducibility to explore containerization and the possibility to just wrap all the dependencies in one piece of software and make sure that even if things change dramatically in the infrastructure, they still have one blob that can be run at any moment, be it in five years, ten years, and that they also can share easily with their colleagues.

Emily: Do you think that that's been accomplished? Has it made it easier for the physicists who collaborate?

Ricardo: I think, yeah, from the experience that I have working with some groups in this area, I think that that has been accomplished. And actually last year, we did some exercise that had two goals: one was to make sure that we could scale out using these technologies to a large amount of resources, and the second was that we could prove that this technology can be used for reusability and reproducibility. And we took, as an example, the analysis that originated the Nobel Prize in 2013, which was the analysis that found the Higgs boson at the time. So, this was code that was pretty old, and the run for this analysis was done in 2012, with old hardware, old code, and what we did is we wrapped the code as it was at the time in a container, and we could run it in infrastructure that is modern today using Kubernetes, and containers, and we could reproduce the results in a much faster way because we have better resources and all this new technologies, but using exactly the same code that was used at the time.

Emily: Wow. And how would you say that moving to Kubernetes and to containers has changed the experience for end-users for the scientists?

Ricardo: Yeah. So, there, I think it depends on which—how to put it—which step of the transition you're at, I think. There's clearly—like, the people that are familiar with Kubernetes, and at ease with these kind of technologies, they will tell about the benefits quite easily. But one thing that we saw is that unlike the transition from physical to virtualization where the resources still felt the same—you would still SSH to your machines, and the tools inside the virtual machines were still very similar—an, like, that—the transition to containerization and to Kubernetes has a steep learning curve. So, we have both some resistance from some people and also some frustration at the start. So, we promote all these trainings, and we do webinars internally to try to help with this, this jump.

But then the measure of benefits once you start using this are quite clear. Like if you're managing a service, you see that deploying updates becomes much much easier. And then if you have to monitor, restart, collect metrics, all of this is well integrated because of this way of declaring how applications should look like, and how they should work, and just delegating to the orchestrator on how to handle this. I think yeah, people have seen this benefit quite a lot from the infrastructure part, from the service management as well.

Emily: What do you think has been surprising or unexpected in the transition to Kubernetes and containers? I think the first thing that I'm interested in is unpleasant surprises: things that were harder than expected.

Ricardo: Yep. I think this steep learning curve is one. Yeah, when you start using containers, it's not that complicated. Like if you—I don’t know—build a Docker file, create your container is not that complicated, but when you jump in the full system, we do see some people having some kind of resistance, I would say, at the start. And it's quite important to have this kind of dissemination activities internally.

And then the other thing that we realized—but it's kind of specific to our infrastructure because I mentioned that we have all these needs for computing capacity, and because we have a large data center, but it's not enough, we actually have a network of institutes around the world and data centers that collaborate with us, and that we developed this technology over the last almost 20 years now. And we had our own ways to, for example, distribute software to these sites. And by moving to containers, this poses a challenge because then the software now is wrapped in this well—defined image; it's distributed by registries, and this doesn't really match all the systems we had before. So, we have to rethink all these aspects of a very large, distributed infrastructure that has limitations, but we know about them and we know how to deal with them, and really learn how to do this when we start using containers, how to distribute the software, how to distribute the images and all the tracking that is required around this. So, for us in this area, it's not something we thought at the start, but it became pretty obvious after a while.

Emily: Makes sense. What specific resistance did you meet, and how did you overcome it?

Ricardo: Yeah, so I think we really need to show the benefit in some cases. We have people that are very proactive, of course, into exploring these technologies, but in some cases—because for many people, this is not their core work. So, if you're a physicist, you want to execute your analysis or get your workloads done. You don't necessarily want to know all the details about infrastructure, so you need to hide this. But if you are involved in this, you kind of have to relearn a bit, how to do system operations and service, how you do your tasks of service manager.

So, you really need a learning process to understand how containers work, how you deploy things in Kubernetes. You have to redefine a lot of the structure of how your application is deployed. If you had a configuration management system that you were using—Puppet, Ansible, whatever—you have to transition to this new model of deploying things. So, I think if you don't see this as part of your core work, you are less motivated to do the transition. But yeah, I think this happens. It happened in the past already, like when we transition to automated configuration management, when we change configuration management systems, there’s always a change that is required.

Yeah, we are also lucky that in this research area, people are open-minded. So, yeah, we do organize a lot of different activities in-house. We do what we call ‘container office hours’ weekly, where people can come and ask questions. We will organize regular webinars in different topics, and we have a lot of internal communication systems where we try to be responsive.

Emily: What do you think are the benefits that, when you encounter somebody who's not super excited about this—or, when you encountered—what are the benefits that you usually highlight?

Ricardo: Yeah. So, actually, we've been collecting from some users that have finished this transition what do they think, and the main benefits they say is in the speed of provisioning. In our previous setup, sometimes you can think that provisioning a cluster can take up to an hour, even in some cases more, to provision a full cluster, while now provisioning a full Kubernetes cluster, even a large one can take less than 10 minutes. And then even more when you do the deployment of the applications or when you do an update to an application, because of the way things are structured, you can really trust that this will happen within seconds and that you can apply these changes very quickly. If you have a large application that is distributed in many nodes, even if you have an automation system, this can take several minutes, or up to an hour or more, for that change to be applied everywhere, while with this containerization and orchestration offered by Kubernetes we can reduce that to just a couple of seconds and be sure that the state is correct; we don't have to do some extra modifications; we just trust the system in most cases. I think those have been the two main aspects.

And then for more advanced use cases, I would say that one benefit is that in some cases, we have tools that are deployed at CERN, but they are also deployed in other sites—as I mentioned, we have this large network—and one thing, because of the how popular Kubernetes became and these technologies, is that they can develop the application at CERN, rapid, and make sure that it runs well on our Kubernetes system, and in most cases, they can just take that, deploy in a different site or in a public cloud, and with no or very little modifications, they will get a similar setup running there. So, that's also a key aspect, is that we reduced the amount of work that is going into custom deployments here and there.

Emily: And would that make it easier for a scientist to, say, collaborate with somebody at a different location?

Ricardo: Yeah, I think that's also one aspect, is that the systems start feeling quite familiar because they're pretty similar in all these places. This is something that is happening in our own infrastructure, but if we look at public clouds, all of them offer some kind of managed Kubernetes service that in some cases have slight differences but deal—in its overall functionality, they're pretty similar. They also have the right abstractions on computing, and also on storage and network so that you don't have to know, necessarily, the details underneath.

Emily: And does this also translate to physicists being able to complete their research faster?

Ricardo: Yeah. So, that's another aspect that we've been investigating is that on-premise, we have this capability to improve our resource usage. So, hopefully, we can do more with the same amount of hardware, but what we've seen is that we can also explore more on-demand resources. So, I mentioned earlier this trial that we did with this Higgs analysis, and the original analysis takes almost a day. We containerized and tried to wrap it all, and we managed to run a subset of the analysis on our infrastructure—but of course, it's a production infrastructure, so we can just ask for all the 300,000 cores to run a test; we get a very small fraction of that if we want to do something like this—but what we showed is that by going, for example, to the public cloud, because the systems are offered in the same way, we took the exact same application, we deployed it in a public cloud, in this case into GCP, using GKE, and we can trust, or we can imagine that they're offering us infinite capacity as long as we pay, but we only pay for what we use.

And basically, we could scale-out, and in this experiment, we scaled out to 25,000 cores. And we could reduce what took initially one day, and then eight hours, we reduced it to something less than six minutes. And that's all also all we pay. Like when we finished the analysis, we just get rid of the resources and stop paying. So, this ability to scale out in this way, explore more resources in different places—we do the same with the HPC computers as well—kind of opens up a lot of possibilities for our physicists to still rely on our large capacity, but explore this possibility of getting extra capacity and do their analysis much faster if they have access to this extra resources.

Emily: Do you ever feel like there's miscommunications or mistranslations that happen between your team—the cloud team, and the end-users—the physicists?

Ricardo: Well, I think their main goal is to get access to resources, so miscommunications usually happen, when the expectations are not met somehow. And this is usually because we have a large amount of users, and we have to prioritize resources. So, sometimes, yeah, sometimes they will want their resources a bit faster than they can get. So, hopefully, this ability to explore more resources from other places will also help in this area.

Emily: Since starting the Cloud transition, what do you think has been easier than expected, either for your team or for the scientists that are using them?

Ricardo: Well, we've done the transition to Cloud itself already back in 2013, offering this API based provisioning of resources. I think one benefit that we had was that in the initial transition to the Cloud, we rely on a software stack of OpenStack for the private cloud, and also, we now rely on Kubernetes and other products that are part of CNCF. For the containerization part, I think one big thing here has been that we always got very involved with the upstream communities, and this has paid off quite a lot. We also contribute a lot to the upstream projects. I think that's the key aspect is that we suddenly realized that things that we used to have to build in-house because we always had big data, and we had to have systems that process all this amount of data, but as big data becomes a more generic problem for other industries and companies, these communities have been built to handle this, and this has been one key aspect that we realize now that we are not alone. We can just join these communities, participate, and contribute to them, learn with them, and we give some but we get a lot more back from these communities.

Emily: What are you looking for from these communities now? And what I'm trying to ask is what's the next step, the problems that you haven't solved yet that you're actively working on?

Ricardo: Yeah, so our main problem right now—well, it's not a problem. It's a challenge I would say—it's that even if we have a lot of data, we are actually having a big upgrade of our experiment in just a few years, where the amount of data could well be increased by something like 100 times, so this gets us to amounts of data that are again pushing the boundaries, and we need to figure out how to handle that. But the big challenge right now is to really understand how to have hybrid infrastructure where our workloads—what we try to do is hide this fact that the infrastructure is maybe not all in our data center, but some bits are in public clouds or other centers, we try to hide this to our user. So, this hybrid approach with the clusters spanning multiple providers is something that still has open issues in terms of networking, but also because we depend so much on this amount of data, we need to understand how to distribute this data in an efficient way and make sure that the workloads run in this places correctly. If you've used public clouds before, you realize quickly that it's very easy to get things up and running, it's very easy to push data, but then when you want to have a more distributed infrastructure and move data around, there's a lot of challenges, both technically and financially. So, these are challenges that we'll have to tackle in the next couple of months or years.

Emily: Yeah, in fact, I was going to ask about how you manage data transfers. I know that could be expensive, and just difficult to manage. Is that something that you feel like you've solved or…?

Ricardo: Yeah, so, technically we know how to do it because we have this past experience with large distributed infrastructures, like this large collection of centers around the world that I mentioned earlier. We have dedicated links between all the sites, and we know how to manage this, and how to efficiently move the data around. And we also are quite well connected to big hubs, to the scientific network hubs in Europe like [JANET], but also direct links to some cloud providers. So, technically we've experimented this, we can have efficient movement of the data we have built over the years to systems that can handle this, and we have the physical infrastructure as well. I think the challenge is understanding the cost models of this: when things are worth moving around, or just move the workloads to where the data is, and how to tackle this. So, we need to understand better how we do this, and also how we account our users when this is happening. This is not a problem that we’ve solved yet.

Emily: Do you also have to have any way to put guardrails on what scientists can and cannot do? To what extent do you need to manage centrally and then let the scientists do whatever they want, but within a certain parameter?

Ricardo: Yeah, so our scientists are grouped in what we call experiments, which are linked to physical experiments underground. These groups are well defined, and the amount of resources that they get are constantly renegotiated, but they are well defined. So, we know which share go to each group of users, but one very important point for us is that we need to try to optimize resource usage. So, if for some reason, a group that has a certain amount of resources allocated is not using them, we allow other groups to just increase a bit—temporarily—their usage, but then we need this mechanism of fair share to balance the usage over a large amount of time. So, using this model on-premises, we know how to do the accounting and management of this. Integrating this to a more flexible infrastructure where we can explore resources from different cloud providers and centers, yeah, it's something we are working on. But we do have this requirement that all of it has to be accounted, and over time, we have to make sure that everyone gets their share.

Emily: Do you think everybody feels like the goals that you set out in moving to containers and Kubernetes, that you've met those goals? I mean, do you think everybody’s, sort of, satisfied that it was a success?

Ricardo: Yeah, I'm not sure we are at that point, yet. I think when we present these exercises, like, what we did that very large scale, that there is a clear improvement. I think people get excited about this. Now, we need to make sure that all of this is offered in a very easy way so that people are happy. But what it also shows—this kind of experiments, and I think it's where physicists get more interested, also, in this kind of technology, is that it makes it much easier to explore not only generic computing resources, but exploring dedicated resources such as accelerators, GPUs, or more custom accelerators like TPUs and IPUs, that we wouldn't have access, necessarily, in our data center, when we offer the potential to have a much easier way to explore these resources and have some sort of central accounting, then, yeah, they get excited because they know that this on-demand offering of GPUs has a lot of potential to speed up their analysis, and without having to overprovision our internal resources, which is something that we wouldn't be allowed to do.

Emily: I'm imagining you don't really have much of a comparison, but I'm thinking about what's unique about being a public sector organization using this technology or being part of the open-source community and what's different from being in a private company.

Ricardo: Yeah, I think one aspect is that we can't necessarily easily choose our providers because it's public sector, so maybe the tendering process is a bit different from a private company. I don’t know exactly, but I could imagine that there are differences in this area. But also the type of workloads we do are quite different. One area that is kind of appearing from industry that is similar to what we do is machine learning, because there's this large batch workloads for machine learning model training. This is a lot of what we do, the majority of what we do.

So, a traditional IT company in the industry would be more interested in providing services. In our case, we do that as well, but I would say 80 percent of our computing is not that. It's large batch submissions that need to execute as fast as possible, so the model is very different. This is especially important when we start talking about multi-cluster because our requirements are much simpler for multi-cluster—or multi-cloud then the requirements of a company that needs to expose services to their end-users.

Emily: Is there anything else that you'd like to add that I haven't thought to ask?

Ricardo: No, I think we covered most of the motivations behind our transition. I would say that I mentioned that one of the challenges is the amount of data that we will have in this upgrade. But I think we did—it could now happen, but it looks like it might happen—we did is that we need to find new computing models to do our analysis than the traditional way we used to do, and this seems to involve, these days, a lot of machine learning. So, I think one thing that we will see that we will benefit a lot, also, from what's happening in the industry is all these frameworks for improving machine learning. And for us, it's also nice to see that a lot of these frameworks are actually relying on this cloud-native Kubernetes type of technologies because it will mean that we are doing something right on the infrastructure, and we might benefit also to upper layers very soon.

Emily: All right, one last question. What is an engineering tool that you can't live without, your favorite tool?

Ricardo: Ah, okay. That’s… well, I would say it's my desktop setup, I would say. I've been a longtime Linux user, and the layout I have, and the tools I use daily are the ones that make me productive, be it to manage my email, calendar, whatever. I think that's one. Yeah, I've learned to rely on Docker for quick prototyping of different environments as well, testing different setups. Yeah, I haven't thought about that. But I would say something in this area.

Emily: And where could listeners connect with you or follow you?

Ricardo: They can contact me by my CERN email which is ricardo.rocha@cern.ch. Or I also have a Twitter account, which is at @ahcorporto. I can pass the name after. But yeah, those are probably the two main points of contact.

Emily: Okay, fabulous. Well, thank you so much for joining me.

Ricardo: Thank you. This was great. Thanks.

Emily: Thanks for listening. I hope you’ve learned just a little bit more about the business of cloud native. If you’d like to connect with me or learn more about my positioning services, look me up on LinkedIn: I’m Emily Omier, that’s O-M-I-E-R, or visit my website which is emilyomier.com. Thank you, and until next time.

Announcer: This has been a HumblePod production. Stay humble.

View Details

Highlights from this episode include:

  • Key market drivers that are causing Cloud Comrade’s clients to containerize applications — including the role that the global pandemic is playing.
  • The pitfalls of approaching cloud migration with a cost-first strategy, and why Andy doesn’t believe in this approach.
  • Common misconceptions that can arise when comparing cloud TCO to on-premise infrastructure.
  • How today’s enterprises tend to view cloud computing versus cloud-native. Andy also mentions a key requirement that companies have to have when integrating cloud services.
  • Andy’s thoughts on build versus buy when integrating cloud services at the enterprise level.
  • Why cloud migration is a relatively safe undertaking for companies because it’s easy to correct mistakes.
  • Why businesses need to re-think AI and to be more realistic in terms of what can actually be automated.
  • Andy’s must-have engineering tool, which may surprise you.

Links:

  • Cloud Comrade LinkedIn: https://www.linkedin.com/company/cloud-comrade/
  • Follow Andy on Twitter: @andywaroma
  • Connect with Andy on LinkedIn: https://www.linkedin.com/in/andyw/

Transcript

Emily: Hi everyone. I’m Emily Omier, your host, and my day job is helping companies position themselves in the cloud-native ecosystem so that their product’s value is obvious to end-users. I started this podcast because organizations embark on the cloud naive journey for business reasons, but in general, the industry doesn’t talk about them. Instead, we talk a lot about technical reasons. I’m hoping that with this podcast, we focus more on the business goals and business motivations that lead organizations to adopt cloud-native and Kubernetes. I hope you’ll join me.

Emily: Welcome to The Business of Cloud Native. I'm Emily Omier, your host, and today I'm here with Andy Waroma. Andy, I just wanted to start with having you introduce yourself.

Andy: Yeah, hi. Thanks, Emily for having me on your podcast. My name is Andy Waroma, and I'm based in Singapore, but originally from Finland. I've been [unintelligible] in Singapore for about 20 years, and for 11 years I spent with a company called SAP focusing on business software applications. And then more recently, about six years ago, I co-founded together with my ex-colleague from SAP, a company called Cloud Comrade, and we have been running Cloud Comrade now for six years and Cloud Comrade focuses on two things: number one, on cloud migrations; and number two, on cloud managed services across the Southeast Asia region.

Emily: What kind of things do you help companies understand when you're helping with cloud migrations? Is this like, like, a lift and shift? To what extent are you helping them change the architecture of their applications?

Andy: Good question. So, typically, if you look at the Southeast Asian market, we are probably anywhere between one to two years behind that of the US market. And I always like to say that the benefit that we have in Southeast Asia is that we have a time machine at our disposal. So, whatever has happened in the US in the past 18 months or so it's going to be happening also in Singapore and Southeast Asia. And for the first three to four years of this business, we saw a lot of lift and shift migrations, but more recently, we have been asked to go and containerize applications to microservices, revamp applications from monolithic approach to a much more flexible and cloud-native approach, and we just see those requirements increasing as companies understand what kind of innovation they can do on different cloud platforms.

Emily: And what do you think is driving, for your clients, this desire to containerize applications?

Andy: Well, if you asked me three months ago, I probably would have said it's about innovation, and business advantage, and getting ahead in the market, and investing in the future. Now, with the global pandemic situation, I would say that most companies are looking at two things: they're looking at cost savings, and they are also looking at automation. And I think cost savings is quite obvious; most companies need to know how they can reduce on their IT expenditure, how they can move from CAPEX to OPEX, how they can be targeting their resources up and down depending on the business demand what they have. And at the same time, they're also not looking to hire a lot of new people into their internal IT organization. So, therefore, most of our customers want to see their applications to be as automated as possible. And of course, microservices, CI/CD pipelines, and everything else helps them to achieve that somewhat. But first and foremost, of course, it's about all services that Cloud provides in general. And then once they have been moving some of those applications and getting positive experiences, that's where we typically see the phase two kicking in, going into cloud-native microservices, containers, Kubernetes, Docker, and so forth.

Emily: And do you think when companies are going into this, thinking, “Oh, I'm going to really reduce my costs.” Do you think they're generally successful?

Andy: I don't think in a way that they think they are. So, especially if I'm looking at the Southeast Asian markets: Singapore, Malaysia, Thailand, Philippines, Indonesia, and perhaps other countries like Vietnam, Myanmar, and Cambodia, it’s a very cost-conscious market, and I always, also like to say that when we go into a meeting, the first question that we get from the customers, “How much?” It is not even what are we going to be delivering, but how much it's going to cost them. That's the first gate of assessment. So, it's very much of an on-premise versus clouds comparison in the beginning.

And I think if companies go in with that type of a mindset, that's not necessarily the winning strategy for them. What they will come to know after a while is that, for example, setting up disaster recovery systems on an on-premise environment, especially when a separate location is extremely expensive, and doing something like that on the Cloud is going to be very cost-efficient. And that's when they start seeing cost savings. But typically, what they will start seeing on Cloud is a process cost-saving, so how they can do things faster, quicker, and be more flexible in terms of responding to end-user demands.

Emily: At the beginning of the process, how much do you think your customers generally understand about how different the cost structure is going to be?

Andy: So, we have more than 200 customers, and we have done more than 500 projects over the six years, and there's a vast range of customers. We have done work with companies with a few people; we have done companies with Fortune 10 organizations, and everything in between, in all kinds of different industries: manufacturing, finance, insurance, public sector, industrial level things, nonprofits, research organizations. So, we can't really say that each customer are same. There are customers who are very sophisticated and they know exactly what they want when going to a cloud platform, but then there are, of course, many other customers who need to be advised much more in the beginning, and that’s where we typically have tools and processes in place where we can assess what is the right proposed strategy for our customers. And then we discuss these ideas in our workshops, and then let the customer decide what will be the best way for them forward, [unintelligible] cost, timelines, and then also potential product risks.

Emily: Are there any common themes around things that companies tend to not understand correctly, or misconceptions that let's say, just come up often; maybe not all the time, but frequently?

Andy: Yeah. I think the most common misconception is that we work a lot with AWS, Microsoft Azure, Google Cloud, Alibaba Cloud, and other cloud platform providers. And typically, when customers are looking to migrate across, they might have virtualized their on-premise environment, and they're looking at the virtual machine to virtual machine cost comparison: how much does it cost them to run one virtual machine on on-premise versus running it on clouds. And when they do that, usually the TCO calculations, they will start breaking down the cost the customer's not taking into consideration all the security that they get with these cloud platforms, the network aspects of it, the compliance aspects, like for example, compliance certifications, regulatory aspects, and all of these things. So, if the customer was to a total TCO of what they're getting on cloud versus what they're getting on their on-premise environment, then they would actually see the huge cost-saving benefits that they are getting. But, at the moment, many of the customers just look at virtual machine to virtual machine comparison, which in our opinion, is wrong.

Emily: So, it can be really hard for companies to do this, sort of, apples to apples comparison because they're not taking everything into account?

Andy: Yeah, to be fair to them, many times you only start getting some of those benefits, of Cloud if you completely decommission your on-premise data center, but if you still the on-premise data center costs and you have to cloud costs, then you can't fully exploit the benefits of Cloud.

Emily: Do you see any common misconceptions related to how you would have answered three months ago, where related to innovation, or agility?

Andy: I think we have seen a huge influx of projects coming to us right now, and I see that a lot of companies who have been sitting on the fence in the previous years, they have now realized that the economic situation is changing very, very rapidly. And it doesn't matter if they have been the market leader, or one of the challengers, they need to do something. And previously, many companies felt that the best approach would be to go to, maybe, one of the large consulting providers, and just get something from them regardless of the price, or alternatively, go and shop around with 10, 15, 20 different vendors and try and squeeze the best price out of them, and [unintelligible] the implementation at the lowest possible cost. I think now, with the pandemic and economic situation, I feel that many companies have been rationalizing their IT strategy, their strategy they’re suspending, and what they are doing is that they are trying to find what is the most suitable vendor given the current situation, forego lengthy RFP processes, do something quickly on Cloud, do some kind of TOC, if it looks good for them, then let's just get going, don't waste internal resources, don't waste external resources, and try to achieve something quickly.

Emily: Do you feel like, in general, your clients tend to understand the difference between just cloud computing and cloud-native?

Andy: Um, I think that definition is probably hazy. I think it’s probably also maybe somehow hazy for me, [laughs], and I think in some cases, the cloud platform providers want to keep it that way as well. Typically, the customers that we are dealing with is relatively traditional enterprises, and we always advise them that rather than doing a lot of transformation on on-premise first, why don't we just help you to lift and shift and move you to cloud, and once you're on Cloud, you can start doing all kinds of experimentation there with cloud-native services. And it's not just about moving these companies to cloud-native services, which have certain benefits, like for example, serverless technologies that reduces the costs a lot, it also requires the companies to have a pretty clear understanding of their internal processes, and then also, critically, an understanding of the roles and responsibilities within those organizations because when you start deploying software and applications in a different way onto a cloud platform, the technology is one part of it, but the process and people is the other part of it, and if the organizations internally are not ready to change the process, they won't be able to reap the benefits of cloud-native services, otherwise.

Emily: And how challenging do you think it is for companies to do the sort of, organizational work to match this sort of change in the way they work together with a different technology?

Andy: I think it is very challenging, and we see a big skill gap there. Many companies, even large companies, they might have issues of attracting talent. And typically, one of the reasons is that even if the companies are wanting to do something on Cloud, they typically choose one Cloud Platform only, and many IT people who are wanting to get into this space or wanting to be in cloud computing. They don't want to focus just on one cloud platform. They want to be focused on multiple different platforms, and they want to understand what are the differences between, let's say, AWS, Azure, Google Cloud, and other providers, and that's why people who are really skilled, they probably end up in consulting companies like ours and other companies. And that's, I believe, is the whole reason of our existence is that many companies cannot do these things in-house, and therefore, they would like to outsource the initial work, and then also the ongoing work terms of managed services to the companies like us.

Emily: Do companies often ask you, “Should we build this internally? Should we outsource this for somebody else to do? Should we buy a tool that does this automatically?” Is that a question that they have, and how do you sort of evaluate an answer?

Andy: [unintelligible] this depends a little bit on the industry. So, for example, there are a number of so-called unicorns, in the region, and these companies have very big IT departments, in-house development department, in-house, they might have thousands of developers, and they feel that the best way for those companies to create differentiation in the market is to build by themselves. But we deal also with a lot of, like I mentioned before, with traditional manufacturing companies, for instance, and there they will go for packaged applications such as SAP, Oracle, .NET applications, and I think that's the best approach for them, rather than start building something in-house, which would not necessarily make sense for them, given the fact that these costs are not in IT industry, but they are in, let’s say, in food and beverages industry. But they've certainly built, and like to build some services on top of those traditional software packages, and analytics and analytical reports, for example, predictive forecasting, and that's something that will bring about uniqueness to their organization, to their process, and competitive advantage. And that are mostly things that they're exploring on the Cloud, and weighing those options for that make sense to buy something prebuilt, or develop it in-house, or outsource it to some organization.

Emily: What do you think your clients tend to be pleasantly surprised about? What ends up being a lot easier than they expect?

Andy: First of all, IT has always been notorious for failed projects. There has been too many of those, and if you look at Cloud implementations, we really haven't had any. I think Cloud is safe in the sense that even if you have estimated, these are the compute requirements, or the storage requirements, or whatever it might be at the initial phase, and it turns out that assessment is wrong, then it's very easy to actually rectify that, and change the services, and if we brought up, let’s say, an SAP on a certain types of instances and two days later, we find out that no, this instance was too big or too small, we can easily go and change that. That also, if I'm taking that SAP example, SAPs customers are typically wanting to move to SAP HANA at some point of the time, but they're unsure whether that makes sense for them from the organizational point of view and business point of view, and we tell them that why don't you just move to Cloud, and then we can provision new instances with SAP HANA. We can migrate your test environment to an SAP HANA environment, and then you can give it a go. So, Cloud allows you to do a lot of this experimentation, which then leads back to the pleasant surprise that there aren't really any failed IP projects, or even if they are failures, the failures are quite minute compared to some of the IT failures they might have seen in the past.

Emily: Interesting. That's a pretty big upside.

Andy: Definitely it is, and we have had some tough, difficult implementations in the past, big migration projects, but with the [unintelligible] from the customer side, and also internal resources, we have always been able to overcome them. And then, even if things are heated at some point of time, usually at the end of the day, the final project sign-off meetings are relatively pleasant and joyous.

Emily: Do you see a major difference in strategy by industry? So, a technically sophisticated company versus a traditional manufacturing company, are they going to follow dramatically different strategies to build competitive advantage using technology?

Andy: Mmm. So, we are really focused on the infrastructure, so when it comes to infrastructure, mine services, [security mine] services, and I'm looking at the infrastructure layer, typically the customers either just choosing to use something on Linux was choosing to use something on Windows. Maybe they use some server technologies and others, but from that point of view, you can't really tell, if you're looking to a customer's cloud console, whether they are a manufacturing company, or whether they are a finance organization. So, in that sense, there is no difference. Where the differences come in, though, is that we, for example, look at financial industries like banks, and insurance companies, that's where the governmental rules, and regulations, and compliance comes in, and there's a lot of paperwork outside of the cloud implementations that need to be done. There's a lot of guidelines that we need to follow, how to provision the infrastructure to be compliant to, let’s say, banking or insurance regulations. That's one industry which is different from the others.

Healthcare in the US, for example, there is HIPAA compliance, so how HIPAA compliant applications need to be built all the way from ground up, so that's also a significant difference. Third one is public sector. And so, like I mentioned before, almost half of our business is coming from public sector, and that's where, for instance, security, and compliance is extremely important, and it sometimes almost goes overboard to a certain extent, but I can understand why. And the security and networking requirements are extreme, and that’s probably, for me, the most different industry out of all of them, customer-based that we have.

And then the last one is multinational companies. Multinational companies have to adhere to their own IT policies, and audit requirements and they also have, many times, a lot of requirements that are relatively unique that they need to follow. Those are also different types of projects from the rest of them. Whether it's a media company, whether it's a manufacturing company, whether it's a research company or organization, I don't see too much differences in the way that they put the demands, and requirements on the infrastructure platform.

Emily: And what do you say is something that your clients sometimes come to you looking for help with, that you still, kind of, don't have a great answer for?

Andy: Machine learning. AI. A lot of companies talk about it, a lot of companies have been pitched about it, a lot of companies come to us, and said we want to do a machine learning project, and I said, “Well, fine. But what do you want to get out of those machine learning algorithms, and what do you want to build?” And essentially, once you start drilling down, what they want to have is this one red button when they press, and everything else can happen magically underneath to resolve all of their business problems.

And I think all of us know that that's not how machine learning works. That's not how AI works, or at least it doesn't work like that right now. So, the expectations that vendors are setting what can be done with machine learning and AI, I think need to be revised. And also, customers need to be a bit more realistic in terms of the expectations what can be automated, what can be done through these new technologies versus what they've been doing in the past.

Emily: Basically, if you want machine learning, you have to know it won't answer all of your questions. You have to know what you want answered, otherwise, you're not even asking the right questions.

Andy: Pretty much. And also, machine learning, just like, also, analytics is that you require vast data sets that need to be ready from the organization, and the data sets need to be relatively clean as well, otherwise, it's hard to come up with models that are accurate, and the old fashioned saying of garbage in, garbage out is also very true, not just in analytics space, but also in the machine learning space.

Emily: Is there anything else that you'd like to add about the topic of cloud-native and what companies get out of it in a business sense?

Andy: I believe that we are living right now in 2020 in a very unique time, and at some point of time, when we start looking back, maybe in a few years from now, we’ll start seeing that actually cloud-native started rocketing and taking off like no one expected in this year, is right now, it's happening. Every company is focusing on it one way or the other, and I think those companies that are taking advantage in this difficult economic times will be the winners once we come out of this economic and pandemic situation.

Emily: I've actually heard that sentiment from many people, so that makes me think it's probably correct. Oh, my last question for you—I almost forgot—what is your can't-live-without engineering tool? So, is there a tool that you would find it difficult to impossible to do your job without?

Andy: It still, after 30 years, its email. So, from that point of view it’s, you know—but if I'm looking at our team, and what they're really focusing on right now, so working on things, like for example, Terraform, or an AWS site on CloudFormation tools, and that brings about automation. And the difficulty, sometimes, in Cloud is that there's so many different things that you can do. So, customers have a perception, when they go to Cloud, they do a one-off project, and then they're kind of done and dusted, and off they go and think about the next thing, but there’s new services coming in all the time, and there is drift that happens on the cloud platform in terms of security, and ports, and network settings, and instances that gets loaded up and everything else, so you can only manage the cloud platform if you manage the [unintelligible] to automation. And I think that some of the [conflict] tools right now for us are starting to become tools like Terraform, from HashiCorp, CloudFormation from AWS, and similar kind of services from Google, and Microsoft that allows us to keep the platform in the kind of a compliance that the customer originally intended the platform to be.

Emily: Where could listeners connect with you, or follow you?

Andy: We have a pretty strong presence on LinkedIn. So, if you just search for Cloud Comrade on LinkedIn. So, that's where we post news about our industry, and our company, and our people, and our customers on a daily basis, and I would be looking forward to also connecting with your listeners.

Emily: Well, thank you so much, Andy. That was very interesting to learn about what you do, and the types of companies that you work with, and what the Southeast Asian market, kind of, looks like.

Andy: Thank you so much, Emily, for your invitation.

Emily: Thanks for listening. I hope you’ve learned just a little bit more about the business of cloud native. If you’d like to connect with me or learn more about my positioning services, look me up on LinkedIn: I’m Emily Omier, that’s O-M-I-E-R, or visit my website which is emilyomier.com. Thank you, and until next time.

Announcer: This has been a HumblePod production. Stay humble.

View Details

Some highlights of the show include:

  • The company’s cloud native journey, which accelerated with the acquisition of Uswitch.
  • How the company assessed risk prior to their migration, and why they ultimately decided the task was worth the gamble.
  • Uswitch’s transformation into a profitable company resulting from their cloud native migration.
  • The role that multidisciplinary, collaborative teams played in solving problems and moving projects forward. Paul also offers commentary on some of the tensions that resulted between different teams.
  • Key influencing factors that caused the company to adopt containerization and Kubernetes. Paul goes into detail about their migration to Kubernetes, and the problems that it addressed.
  • Paul’s thoughts on management and prioritization as CTO. He also explains his favorite engineering tool, which may come as a surprise.

Links:

  • RVU Website: https://www.rvu.co.uk/
  • Uswitch Website: https://www.uswitch.com/
  • Twitter: https://twitter.com/pingles
  • GitHub: https://github.com/pingles

Transcript

Announcer: Welcome to The Business of Cloud Native podcast, where we explore how end users talk and think about the transition to Kubernetes and cloud-native architectures.

Emily: Welcome to The Business of Cloud Native. I'm your host, Emily Omier, and today I am chatting with Paul Ingles. Paul, thank you so much for joining me.

Paul: Thank you for having me.

Emily: Could you just introduce yourself: where do you work? What do you do? And include, sort of, some specifics. We all have a job title, but it doesn't always reflect what our actual day-to-day is.

Paul: I am the CTO at a company called RVU in London. We run a couple of reasonably big-ish price comparison, aggregator type sites. So, we help consumers figure out and compare prices on broadband products, mobile phones, energy—so in the UK, energy is something which is provided through a bunch of different private companies, so you've got a fair amount of choice on kind of that thing. So, we tried to make it easier and simpler for people to make better decisions on the household choices that they have. I've been there for about 10 years, so I've had a few different roles. So, as CTO now, I sit on the exec team and try to help inform the business and technology strategy. But I've come through a bunch of teams. So, I've worked on some of the early energy price comparison stuff, some data infrastructure work a while ago, and then some underlying DevOps type automation and Kubernetes work a couple of years ago.

Emily: So, when you get in to work in the morning, what types of things are usually on your plate?

Paul: So, I keep a journal. I use bullet journalling quite extensively. So, I try to track everything that I’ve got to keep on top of. Generally, what I would try to do each day is catch up with anybody that I specifically need to follow up with. So, at the start of the week, I make a list of every day, and then I also keep a separate column for just general priorities.

So, things that are particularly important for the week, themes of work going on, like, technology changes, or things that we're trying to launch, et cetera. And then I will prioritize speaking to people based on those things. So, I'll try and make sure that I'm focusing on the most important thing. I do a weekly meeting with the team. So, we have a few directors that look after different aspects of the business, and so we do a weekly meeting to just run through everything that's going on and sharing the problems. We use the three P's model: so, sharing progress problems and plans. And we use that to try and steer on what we do. And we also look at some other team health metrics.

Yeah, it's interesting actually. I think when I switched from working in one of the teams to being in the CTO role, things change quite substantially. That list of things that I had to care about increase hugely, to the point where it far exceeded how much time I had to spend on anything. So, nowadays, I find that I'm much more likely for some things to drop off. And so it's unfortunate, and you can't please everybody, so you just have to say, “I'm really sorry, but this thing is not high on the list of priorities, so I can't spend any time on it this week, but if it's still a problem in a couple of weeks time, then we'll come back to it.” But yeah, it can vary quite a lot.

Emily: Hmm, interesting. I might ask you more questions about that later. For now, let's sort of dive into the cloud-native journey. What made RVU decide that containerization was a good idea and that Kubernetes was a good idea? What were the motivations and who was pushing for it?

Paul: That's a really good question. So, I got involved about 10 years ago. So, I worked for a search marketing startup that was in London called Forward Internet Group, and they acquired USwitch in 2010. And prior to working at Forward, I'd worked as a consultant at ThoughtWorks in London, so I spent a lot of time working in banks on continuous delivery and things like that. And so when Uswitch came along, there were a few issues around the software release process. Although there was a ton of automation, it was still quite slow to actually get releases out. We were only doing a release every fortnight. And we also had a few issues with the scalability of data.

So, it was a monolithic Windows Microsoft stack. So, there was SQL Server databases, and .NET app servers, and things like that. And our traffic can be quite spiky, so when companies are in the news, or there's policy changes and things like that, we would suddenly get an increase in traffic, and the Microsoft solution would just generally kind of fall apart as soon as we hit some kind of threshold. So, I got involved, partly to try and improve some of the automation and release practices because at the search start-up, we were releasing experiments every couple of hours, even.

And so we wanted to try and take a bit of that ethos over to Uswitch, and also to try and solve some of the data scalability and system scalability problems. And when we got started doing that, a lot of it was—so that was in the early heyday of AWS, so this was about 2008, that I was at the search startup. And we were used to using EC2 to try and spin up Hadoop clusters and a few other bits and pieces that we were playing around with. And when we acquired Uswitch, we felt like it was quickest for us to just create a different environment, stick it under the load balancer so end users wouldn't realize that some requests was being served off of the AWS infrastructure instead, and then just gradually go from there. We found that that was just the fastest way to move.

So, I think it was interesting, and it was both a deliberate move, but it was also I think the degree to which we followed through on it, I don't think we'd really anticipated quite how quickly we would shift everything. And so when Forward made the acquisition, I joined summer of 2010, and myself and a colleague wrote a little two-pager on, here are the problems we see, here are the things that we think we can help with and the ways that technology approach that we'd applied at Forward would carry across, and what benefits we thought it would bring. Unfortunately because Forward was a privately held business—we were relatively small but profitable—and the owner of that business was quite risk-affine. He was quite keen on playing blackjack and other stuff. So, he was pretty happy with talking about probabilities of success.

And so we just said, we think there's a future in it if we can get the wheels turning a bit better. And he was up for it. He backed us and we just took it from there. And so we replaced everything from self-hosted physical infrastructure running on top of .NET to all AWS hosted, running a mix of Ruby, and Closure, and other bits and pieces in about two years. And that's just continued from there. So, the move to Kubernetes was a relatively recent one; that was only within the last—I say ‘recent.’ it was about two years ago, we started moving things in earnest. And then you asked what was the rationale for switching to Kubernetes—

Emily: Let me first ask you, when you were talking with the owner, what were the odds that you gave him for success?

Paul: [laughs]. That's a good question. I actually don't know. I think we always knew that there was a big impact to be had. I don't think we knew the scale of the upside. So, I don't think we—I mean, at the time, Uswitch was just about breaking even, so we didn't realize that there was an opportunity to radically change that. I think we underestimated how long it would take to do.

So, I think we’d originally thought that we could replace, I think maybe most of the stuff that we needed replaced within six months. We had an early prototype out within two weeks, two or three weeks because we'd always placed a big emphasis on releasing early, experimenting, iterative delivery, A/B testing, that kind of thing. So, I think it was almost like that middle term that was the harder piece. And there was definitely a point where… I don't know, I think it was this classic situation of pulling on a ball of string where it was like, what wanted to do was to focus on improving the end-user experience because our original belief was that, aside from the scalability issues, that the existing site just didn't solve the problem sufficiently well, that it needed an overhaul to simplify the journeys, and simplify the process, and improve the experience for people.

We were focusing on that and we didn't want to get drawn into replacing a lot of the back office and integration type systems partly because there was a lot of complexity there. But also because you then have to engage with QA environments, and test environments, and sign-offs with the various people that we integrate with. But it was, as I said, it was this kind of tugging on a ball of string where every improvement that we made in the end-user experience—so we would increase conversion rate by 10 percent but through doing that, we would introduce downstream error in the ways that those systems would integrate, and so we gradually just ended up having to pull in slightly more and more pieces to make it work. I don't think we ever gave odds of success. I think we underestimated how long that middle piece would take. I don't think we really anticipated the degree of upside that we would get as a consequence, through nothing other than just making releases quicker, being able to test and move faster, and focusing on end-user experience was definitely the right thing to focus on.

Emily: Do you think though, that everybody perceived it as a risk? I'm just asking because you mentioned the blackjack, was this a risk that could fail?

Paul: Well, I think the interesting thing about it was that we knew it was the right thing to do. So, again, I think our experience as consultants at ThoughtWorks was on applying continuous delivery, what we would today call DevOps, applying those practices to software delivery. And so we'd worked on systems where there weren't continuous integration servers and where people weren't releasing every day, and then we’d worked in environments where we were releasing every couple of hours, and we were very quickly able to hone in on what worked and discard things that didn't. And so I think because we've been able to demonstrate that success within the search business, I think that carried a great deal of trust.

And so when it came to talking about things we could potentially do, we were totally convinced that there were things that we could improve. I think it was a combination of, there was a ton of potential, we knew that there was a new confluence of technologies and approaches that could be successful if we were able to just start over. And then I think also probably a healthy degree of, like, naive, probably overconfidence in what we could do that we would just throw ourselves into it. So, it's hard work, but yeah, it was ultimately highly successful. So, it's something I'm exceedingly proud of today.

Emily: You said something really interesting, which is that Uswitch was barely profitable. And if I understand correctly, that changed for the better. Can you talk about how this is related?

Paul: Yeah, sure. I think the interesting thing about it was that we knew that there was something we could do better, but we weren't sure what it was. And so the focus was always on being able to release as frequently as we possibly could to try and understand what that was, as well as trying to just simplify and pay back some of the technical debt. Well, so, trying to overcome some of the artificial constraints that existed because of the technology choices that people have made—perfectly decent decisions on, back in the day, but platforms like AWS offered better alternatives, now. So, we just focused on being able to deliver iteratively, and just keep focusing on continual improvement, releasing, understanding what the problems were, and then getting rid of those little niggly things.

The manager I had at Forward was this super—I don't know, he just had the perfect ethos, and he was driven—so we were a team that were focused on doing daily experiments. And so we would rely on data on our spend and data on our revenue. And that would come in on a daily cycle. So, a lot of the rhythm of the team was driven off of that cycle. And so as we could run experiments and measure their profitability, we could then inform what we would do on the day.

And so, we have a handful of long-running technology things that we were doing, and then we would also have other tactical things that he would have ideas on, he would have some hypothesis of, well, “Maybe this is the reason that this is happening, let's come up with a test that we can use to try and figure out whether that's true.” We would build something quickly to throw it together to help us either disprove it or support it, and we would put it live, see what happened, and then move on to the next thing. And so I think a lot of the—what we wanted to do is to instill a bit of that environment in Uswitch. And so a lot of it was being able to release quickly, making sure that people had good data in front of them. I mean, even tools like Google Analytics were something which we were quite au fait with using but didn't have broad adoption at the time. And so we were using that to look at site behavior and what was going on and reason about what was happening. So, we just tried to make sure that people were directly using that, rather than just making changes on a longer cycle without data at all.

Emily: And can you describe how you were working with the business side, and how you were communicating, what the sort of working relationship was like? If there was any misunderstandings on either side.

Paul: Yeah, it’s a good question. So, when I started at Uswitch, the organizational structure was, I guess, relatively classical. So, you had a pooled engineering team. So, it was a monolithic system, deployed onto physical infrastructure. So, there was an engineering team, there was an operations team, and then there were a handful of people that were business specific in the different markets that we operated in. So, there was a couple of people that focused on, like, the credit card market; a couple of people that focused on energy, for example.

And, I used to call it the stand-up swarm: so, in the morning, we would sit on our desks and you would see almost the entire office moved from the different card walls that were based around the office. Although there was a high degree of interaction between the business stakeholders, the engineers, designers, and other people, it always felt slightly weird that you would have almost all of the company interested in almost everything that was going on, and so I think the intuition we had was that a lot of the ways that we would think about structuring software around loosely-coupled but highly cohesive, those same principles should or could apply to the organization itself. And so what we tried to do is to make sure that we had multidisciplinary teams that had the people in them to do the work. So, for the early days of the energy work, there was only a couple of us that were in it.

So, we had a couple of engineers, and we had a lady called Emma, who was the product owner. She used to work in the production operations team, so she used to be focused on data entry from the products that different energy providers would send us, but she had the strongest insight into the domain problem, what problem consumers were trying to overcome, and what ways that we could react to it. And so, when we got involved, she had a couple of ideas that she'd been trying to get traction on, that she'd been unable to. And so what we—we had a, I don't know, probably a, I think a half-day session in an office. So, we took over the boardroom at the office and just said, “Look, we could really do with a separate space away from everybody to be able to focus on it. And we just want to prove something out for a couple of weeks. And we want to make sure that we've got space for people to focus.”

And so we had a half-day in there, we had a conversation about, “Okay, well, what's the problem? What's the technical complexity of going after any of these things?” And there's a few nuances, too. Like, if you choose option A, then we have to get all of the historical information around it, as well as the current products and market. Whereas if we choose option B, then we can simplify it down, and we don't need to do all of that work, and we can try and experiment with something sooner.

So, we wanted it to be as collaborative as possible because we knew that the way that we would be successful is by trying to execute on ideas faster than we’d been able to before. And at the same time, we also wanted to make sure that there was a feeling of momentum and that we would—I think there was probably a healthy degree of slight overconfidence, but we were also very keen to be able to show off what we could do. And so we genuinely wanted to try and improve the environment for people so that we could focus on solving problems quicker, trying out more experiments, being less hung up on whether it was absolutely the right thing to do, and instead just focus on testing it. So, were there tensions? I think there were definitely tensions; I don't think there weren’t tensions so much on the technical side; we were very lucky that most of the engineers that already worked there were quite keen on doing something different, and so we would have conversations with them and just say, “Look, we'll try everything we can to try and remove as many of the constraints that exist today.”

I think a lot of the disagreement or tension was whether or not it was the right problem to be going after. So, again, the search business that we worked in was doing a decent amount of money for the number of people that were there, and we knew there was a problem we could fix, but we didn't know how much runway it would have. And so there was a lot of tension on whether we should be pulling people into focusing on extending the search business, or whether we needed to focus on fixing Uswitch. So, there was a fair amount of back and forth about whether or not we needed to move people from one part of the business to another and that kind of thing.

Emily: Let's talk a little bit about Kubernetes, and how Uswitch decided to use Kubernetes, what problem it solved, and who was behind the decision, who was really making the push.

Paul: Yeah, interesting. So, I think containers was something that we'd been experimenting with for a little while. So, as I think a lot of the culture was, we were quite risk-affine. So, we were quite keen to be trying out new technologies, and we'd been using modern languages and platforms like Closure since the early days of them being available. We’d been playing around with containers for a while, and I think we knew there was something in it, but we weren't quite sure what it was.

So, I think, although we were playing around with it quite early, I think we were quite slow to choose one platform or another. I think in the end, we—in the intervening period, I guess, between when we went from the more classical way of running Puppet across a bunch of EC2 instances that run a version of your application, the next step after that was switching over to using ECS. So, Amazon's container service. And I guess the thing that prompted a bit more curiosity into Kubernetes was that—I forgot the projects I was working on, but I was working on a team for a little while, and then I switched to go do something else. And I needed to put a new service up, and rather than just doing the thing that I knew, I thought, “Well, I'll go talk to the other teams.” I'll talk to some other people around the company, and find out what's the way that I ought to be doing this today, and there was a lot of work around standardizing the way that you would stand up an ECS cluster.

But I think even then, it always felt like we were sharing things in the wrong way. So, if you were working on a team, you had to understand a great deal of Amazon to be able to make progress. And so, back when I got started at Uswitch, when I talk about doing the work about the energy migration, AWS at the time really only offered EC2, load balancers, firewalling, and then eventually relational databases, and so back then the amounts of complexity to stand up something was relatively small. And then come to a couple of years ago. You have to appreciate and understand routing tables, VPCs, the security rules that would permit traffic to flow between those, it was one of those—it was just relatively non-trivial to do something that was so core to what we needed to be able to do.

And I think the thing that prompted Kubernetes was that, on the Kubernetes project side, we'd seen a gradual growth and evolution of the concepts, and abstractions, and APIs that it offered. And so there was a differentiation between ECS or—I actually forget what CoreOS’s equivalent was. I think maybe it was just called CoreOS. But there are a few alternative offerings for running containerized, clustered services, and Kubernetes seems to take a slightly different approach that it was more focused on end-user abstractions. So, you had a notion of making a deployment: that would contain replicas of a container, and you would run multiple instances of your application, and then that would become a service, and you could then expose that via Ingress. So, there was a language that you could use to talk about your application and your system that was available to you in the environment that you're actually using.

Whereas AWS, I think, would take the view that, “Well, we've already got these building blocks, so what we want our users to do is assemble the building blocks that already exist.” So, you still have to understand load balancers, you still have to understand security groups, you have to understand a great deal more at a slightly lower level of abstraction. And I think the thing that seemed exciting, or that seems—the potential about Kubernetes was that if we chose something that offered better concepts, then you could reasonably have a team that would run some kind of underlying platform, and then have teams build upon that platform without having to understand a great deal about what was going on inside. They could focus more on the applications and the systems that they were hoping to build. And that would be slightly harder on the alternative.

So, I think at the time, again, it was one of those fortunate things where I was just coming to the end of another project and was in the fortunate position where I was just looking around at the various different things that we were doing as a business, and what opportunity there was to do something that would help push things on. And Kubernetes was one of those things which a couple of us had been talking about, and thinking, “Oh, maybe now is the time to give it a go. There's enough stability and maturity in it; we're starting to hit the problems that it's designed to address. Maybe there's a bit more appetite to do something different.”

So, I think we just gave it a go. Built a proof of concept, showed that could run the most complex system that we had, and I think also did a couple of early experiments on the ways in which Kubernetes had support for horizontal scaling and other things which were slightly harder to put into practice in AWS. And so we did all that, I think gradually it just kind of growed out from there, just took the proof of concepts to other teams that were building products and services. We found a team that were struggling to keep their systems running because they were a tiny team. They only had, like, two or three engineers in. They had some stability problems over a weekend because the server ran out of hard disk space, and we just said, “Right. Well, look, if you use this, we'll take on that problem. You can just focus on the application.” It kind of just grew and grew from there.

Emily: Was there anything that was a lot harder than you expected? So, I'm looking for surprises as you're adopting Kubernetes.

Paul: Oh, surprises. I think there was a non-trivial amount that we had to learn about running it. And again, I think at the point at which we'd picked it up, it was, kind of, early days for automation, so there was—I think maybe Google had just launched Google Kubernetes engine on Google Cloud. Amazon certainly hadn't even announced that hosted Kubernetes would be an option. There was an early project within Kubernetes, called kops that you could use to create a cluster, but even then it didn't fit our network topology because it wouldn't work with the VPC networking that we needed and expected within our production infrastructure.

So, there was a lot of that kind of work in the early days, to try and make something work, you had to understand in quite a level of detail what each component of Kubernetes was doing. As we were gradually rolling it out, I think the things that were most surprising were that, for a lot of people, it solved a lot of problems that meant they could move on, and I think people were actually slightly surprised by that. Which, [laughs], it sounds like quite a weird turn of phrase, but I think people were positively surprised at the amount of stuff that they didn't have to do for solving a fair few number of problems that they had. There was a couple of teams that were doing things that are slightly larger scale that we had to spend a bit more time on improving the performance of our setup. So, in particular, there was a team that had a reasonably strong requirement on the latency overheads of Ingress.

So, they wanted their application to respond within, I don't know, I think it was maybe 200 milliseconds or something. And we, through setting up the monitoring and other bits and pieces that we had, we realized that Ingress currently was doing all right, but there was a fair amount of additional latency that was added at the tail that was a consequence of a couple of bugs or other things that existed in the infrastructure. So, there was definitely a lot of little niggly things that came up as we were going, but we were always confident that we could overcome it. And, as I said, I think that a lot of teams saw benefits very early on. And I think the other teams that were perhaps a little bit more skeptical because they got their own infrastructure already, they knew how to operate it, it was highly tested, they'd already run capacity and load tests on it, they were convinced that it was the most efficient thing that they could possibly run. I think even over the long run, I think they realized that there was more work that they needed to do than they should be focusing on, and so they were quite happy to ultimately switch over to the shared platform and infrastructure that the cloud infrastructure team run.

Emily: As we wrap up, there's actually a question I want to go back to, which is how you were talking about the shifting priorities now that you've become CTO. Do you have any sort of examples of, like, what are the top three things that you will always care about, that you will always have the energy to think about? And then I'm curious to have some examples of things that you can't deal with, you can’t think about. The things that tend to drop off.

Paul: The top three things that I always think about. So, I think, actually, what's interesting about being CTO, that I perhaps wasn't expecting is that you're ever so slightly removed from the work, that you can't rely on the same signals or information to be able to make a decision on things, and so when I give the Kubernetes story, it's one of those, like, because I’d moved from one system to another, and I was starting a new project, I experienced some pain. It’s like, “Right. Okay, I've got to go do something to fix this. I've had enough.”

And I think the thing that I'm always paying attention to now, is trying to understand where that pain is next, and trying to make sure that I've got a mechanism for being able to appreciate that. So, I think a lot of the things I try to spend time on are things to help me keep track of what's going on, and then help me make decisions off the back of it. So, I think the things that I always spend time on are generally things trying to optimize some process or invest in automation. So, a good example at the moment is, we're talking about starting to do canary deployments. So, starting to automate the actual rollout of some new release, and being able to automate a comparison against the existing service, looking at latency, or some kind of transactional metrics to understand whether it's performing as well or different than something historical.

So, I think the things that I tend to spend time on are process-oriented or are things to try and help us go quicker. One of the books that I read that changed my opinion of management was Andy Grove’s, High Output Management. And I forget who recommended it to me, but somebody recommended it to me, and it completely altered my opinion of what value a manager can add. So, one of the lenses I try to apply to anything is of everything that's going on, what's the handful of things that are going to have the most impact or leverage across the organization, and try and spend my time on those. I think where it gets tricky is that you have to go broad and deep. So, as much as there are broad things that have a high consequence on the organization as a whole, you also need an appreciation of what's going on in the detail, and I think that's always tricky to manage. I'm sorry, I forgot what was the second part of your question.

Emily: The second part was, do you have any examples of the things that you tend to not care about? That presumably someone is asking you to care about, and you don't?

Paul: [laughs]. Yeah, it’s a good question. I don't think it's that I don't care about it. I think it's that there are some questions that come my way that I know that I can defer, or they're things which are easy to hand off. So, I think the… that is a good question. I think one of the things that I think are always tricky to prioritize, are things which feel high consequence but are potentially also very close to bikeshedding.

And I think that is something which is fair—I'd be interested to hear what other people said. So, a good example is, like, choice of tooling. And so when I was working on a team, or on a problem, we would focus on choosing the right tool for the job, and we would bias towards experimenting with tools early, and figuring out what worked, and I think now you have to view the same thing through a different lens. So, there's a degree to which you also incur an organizational cost as a consequence of having high variability in the programming languages that you choose to use. And so I don't think it's something I don't care about, but I think it's something which is interesting that I think it's something which, over the time I've been doing this role, I've gradually learned to let go of things that I would otherwise have previously thoroughly enjoyed getting involved in.

And so you have to step back and say, “Well, actually I'm not the right person to be making a decision about which technology this team should be using. I should be trusting the team to make that decision.” And you have to kind of—I think that over the time I've been doing the role, you kind of learn which are the decisions that are high consequence that you should be involved in and which are the ones that you have to step back from. And you just have to say, look, I've got two hours of unblocked time this week where I can focus on something, so of the things on my priority list—the things that I've written in my journal that I want to get done this month—which of those things am I going to focus on, and which of the other things can I leave other people to get on with, and trust that things will work out all right?

Emily: That's actually a very good segue into my final question, which is the same for everyone. And that is, what is an engineering tool that you can't live without—your favorite?

Paul: Oh, that’s a good question. So, I don't know if this is a cop-out by not mentioning something engineering-related, but I think the tool and technique which has helped me the most as I had more and more management responsibility and trying to keep track of things, is bullet journaling. So, I think, up until, I don’t know, maybe five years ago, probably, I'd focus on using either iOS apps or note tools in both my laptop, and phone, and so on, and it never really stuck. And bullet journaling, through using a pen and a notepad, it forced me to go a bit slower. So, it forced me to write things down, to think through what was going on, and there is something about it being physical which makes me treat it slightly differently.

So, I think bullet journaling is one of the things which has had the—yeah, it's really helped me deal with keeping track of what's going on, and then giving me the ability to then look back over the week, figure out what were the things that frustrated me, what can I change going into next week, one of the suggestions that the person that came up with bullet journaling recommended, is this idea of an end of week reflection. And so, one of the things I try to do—it's been harder doing it now that I'm working at home—is to spend just 15 minutes at the end of the week thinking of, what are the things that I'm really proud of? What are some good achievements that I should feel really good about going into next week? And so I think a lot of the activities that stem from bullet journaling have been really helpful. Yeah, it feels like a bit of a cop-out because it's not specifically technology related, but bullet journaling is something which has made a big difference to me.

Emily: Not at all. That's totally fair. I think you are the first person who's had a completely non-technological answer, but I think I've had someone answer Slack, something along those lines.

Paul: Yeah, I think what's interesting is there there are loads of those tools that we use all the time. Like Google Docs is something I can't live without. So, I think there's a ton of things that I use day-to-day that are hard to let go off, but I think the I think that the things that have made the most impact on my ability to deal with a stressful job, and give you the ability to manage yourself a little bit, I think yeah, it's been one of the most interesting things I've done.

Emily: And where could listeners connect with you or follow you?

Paul: Cool. So, I am @pingles on Twitter. My DMs are open, so if anybody wants to talk on that, I'm happy to. I’m also on GitHub under pingles, as well. So, @pingles, generally in most places will get you to me.

Emily: Well, thank you so much for joining me.

Paul: Thank you for talking. It's been good fun.

Announcer: Thank you for listening to The Business of Cloud Native podcast. Keep up with the latest on the podcast at thebusinessofcloudnative.com and subscribe on iTunes, Spotify, Google Podcasts, or wherever fine podcasts are distributed. We'll see you next time.

View Details

Some of the highlights include:

  • Why Vodafone moved to a cloud native architecture. As Tom explains, the company was struggling to manage operations across more than 20 markets. They also needed to improve the customer experience, and foster customer loyalty.
  • Why their business and engineering teams were both in favor of cloud native.
  • The benefits of deploying daily operational activities around a single cloud native platform.
  • An overview of where Vodavone currently is in their overall cloud native journey. Tom also explains how cloud native conversations have changed inside of the company throughout their journey, as various business units have caught on to the benefits of the cloud.
  • Vodafaone’s transition from outsourcing roughly 97 percent of their operations, to bringing 95 percent in house. Tom explains how this has improved efficiency and expedited time to market.
  • The challenge that Vodafone faced in trying to apply legacy network security solutions to distributed and dynamic systems.
  • Tom’s thoughts on why Vodafone’s cloud native transition and modernization efforts have been crucial to their success over the last five years.

Links:

  • Vodafone Group: https://www.vodafone.com/
  • Connect with Tom on LinkedIn: https://uk.linkedin.com/in/tom-kivlin-5b469321
  • The Business of Cloud Native: http://thebusinessofcloudnative.com
  • Tom’s Twitter: https://twitter.com/tomkivlin
  • CNCF GitHub: https://github.com/cncf
  • CNCF Slack: https://slack.cncf.io/
  • Kubernetes Slack: http://slack.kubernetes.io/

Transcript

Announcer: Welcome to The Business of Cloud Native podcast, where we explore how end users talk and think about the transition to Kubernetes and cloud-native architectures.

Emily: Welcome to The Business of Cloud Native. I'm Emily Omier, your host, and today I am chatting with Tom Kivlin. Tom, thank you so much for joining us.

Tom: You're welcome. No problem.

Emily: Let's just start out with having you introduce yourself. What do you do? Where do you work, and what do you actually do during your workday?

Tom: Sure. So, I'm a principal cloud orchestration architect at Vodafone Group. I work in the UK. And my day job consists of providing guidance and strategy and architectural blueprints for cloud-native platforms within Vodafone. So, that's around providing guidance to the software domains that are looking to adopt cloud-native architectures and methodologies and also to the more traditional infrastructure domains to try and help them provide their services in a more cloud-native manner to those modern teams.

Emily: And what does that mean when you go into the office—or your home office, go into your dining room where your laptop is, I don't know—what do you actually do? What does an average day look like?

Tom: It can vary. So, depending on the activity at the time, it could be anything from preparing a global policy that needs to go through the senior technology leadership team, to preparing some extremely detailed requirements for selection process or creating some infrastructures code, or the code artifacts for the deployment of cloud-native services, whether that's in our lab, or to help our services teams within Vodafone.

Emily: Tell me a little bit more about what pain made Vodafone think about moving to cloud-native and Kubernetes.

Tom: Primarily, it was the challenge of having 25 different markets, or 23 now. We launched a digital strategy to—so back in 2015, we launched a five-year strategy, which we wanted to massively increase the rollout of 4G, of converged network offerings, of improved customer experience. And we found that the traditional way of managing software was not supportive enough in our ambition. And so, having to choose cloud-native technologies, things like Kubernetes, but also the modern operating models, that was the driver: it was to improve our customer experience, and our customer-affecting KPIs, really.

Emily: And when you say it wasn't supportive enough, what do you mean specifically?

Tom: So, things like time to market, for example. So, if we wanted to offer a new service—so one of the things that 4G started the drive towards was a more granulated service offering to consumers, and so lots of different things could be offered. And if it took you six months to think of an idea and then have to go through—or even longer than six months to get to the point where that could be offered to customers, even if it was just a very minor feature within an existing product, then that's not going to engender customer loyalty. And so, things like the cloud-native mindset, where there's a much closer link between the engineering teams and the customer, there are much shorter periods of time between ideas coming in from the customers and then being delivered back to the customers as product features, that sort of time to market was really enabled by cloud-native technologies and mindsets.

Emily: And how does having two dozen, more or less, different markets, how does that play into the decision A) to move to cloud-native in general, and managing the IT infrastructure?

Tom: So, one of the things that's really driven it is trying to simplify and reuse artifacts. So, if you've got 23 markets all doing a different thing, then there's obviously a lot of duplication happening across the group, whereas if everyone's using the same technology in the same platforms—take Kubernetes as the example—everyone can write their software for that platform. Everyone can write their operational ecosystem around that platform. So, the deployment artifacts, the pipelines, the day two operational activities, they can all be based around that single cloud-native platform. And so, that enables a huge amount of efficiency from the operational side. And that in turn allows those engineering teams to focus on things that are adding value to the business and the customer instead of having to focus on fairly low-level tasks that are just keeping the lights on, if you like.

Emily: What's different for each one of those markets?

Tom: So, it might be something like language, it might be something as simple as that. It may be that the offerings are slightly tweaked. So, rather than, I don't know, as an example, rather than Spotify being included as a kind of add on, it might be some other service that's more relevant to that market. It may be that there are particular regulatory requirements that are specific to a market that needs to be considered within the product design and the engineering of it. And so, having a cloud-native response allows sharing and reuse of artifacts where we can, but still allows for that customization where it's required.

Emily: Where would you say Vodafone is in the cloud-native journey? Do you feel like you've, mission accomplished?

Tom: So, mission accomplished, as in the first step, yeah. So, we set out a goal in 2015, to get a certain number of our applications to the Cloud, and that's largely been reached, I think, especially with our customer channels, so that the kind of points of interaction with the customer, the huge number of those are cloud-native today. And things like automated customer interaction with chatbots, and the like, that's all added to the cloud-nativeness of the interaction. As part of our next iteration, we'll be looking for more cloud-native software and cloud-native platforms, and that will start extending into the network systems themselves, as well as the more digital and easily modernizable layers, if you like.

Emily: What sort of business value do you feel like you're looking for as you move to the next step?

Tom: So, primarily, it's going to be driven by customer satisfaction and customer affecting KPIs, like I said before. That's always what’s driven the business metrics anyway. So, things like being able to support the demand of the customer. So, whether that's the new 5G services, for increased bandwidth. So, obviously, if our network systems themselves are cloud-native, then taking advantage of the auto-scaling, and the auto-healing, and the autonomic nature, then the customer experience, and the customer satisfaction will increase.

Improving time to market, so again, part of 5G is that the whole notion of creating more differentiating services, and so if we can do that through the cloud-native mindset with product owners being much more closely engaged with customers, then that improves our product offerings. And we can optimize our network profitability by using cloud-native features like modern big data analytics, and even AI and automation to improve the operations of the network. At the end of the day, the business value is improved customer satisfaction, which improves our financial performance, obviously.

Emily: And when you started out in 2015, who was pushing for moving to cloud-native? Was this the business saying, “Hey, how do we improve customer satisfaction?” Was it engineering saying, “Hey, here's an idea for something that could help us move faster?” Who was behind that?

Tom: That’s a good question. I think it's probably an element of both. It was the opposite of the push me, pull you, I guess. So, there was engineering pushing on an open door, I suppose you could say. So, Cloud was a bit of a buzzword around that time anyway, but I think it's fair to say the concepts of improved time to market, improved stability, the potential for improved security, improved automation, and repeatability, they were all relatively easy sells to product teams who want to be able to sell products to customers. And once you're able to explain what problems those concepts solve, I think it became a bit of a, like I say, pushing on an open door.

Emily: Can you tell me a little bit about the process of explaining what problems these things solve? Was there anything that was getting lost in translation?

Tom: Yeah. I think the biggest thing that I can recall—obviously, it's a company-wide thing. I'm never going to be aware of everything that happens—certainly, it's critical to try and understand what the target operating model is before trying to say, “Here's the technology solution to it.” So, I think some of the lessons that were learnt in the early stages were, rather than trying to say, “Here's the technology answer to a modern way of working that hasn't been agreed or adopted or even understood yet,” let's do that part first, so people understand how they need to work in this modern kind of culture. And then the technology answers then make a bit more sense to people because they're able to say, “Okay, I understand the problems that’s solving now because I'm now working in that way of working.” So, that's probably the biggest learning point I would take from the previous five years.

Emily: Do you feel like the conversation, how did it evolve from the first conversations over the course of the past five years, and then what's it like now?

Tom: It's very different now. The concept of Cloud and cloud-native has become a given and very well understood across the business, even outside of technology. So, we talked to other business units, and they're quite comfortable in understanding the benefits of Cloud. And it's now about when they mature into cloud-native, and when they mature operating models, rather than if. And it's now talking and giving guidance about how to do it, rather than trying to sell the concept itself. So, it just feels like you're at that next stage of not having to sell the idea anymore, and more into the detail of how to implement that idea.

Emily: What would you say were some of the biggest surprises? And let's start with thinking about some of the biggest surprises, not necessarily technically but organizationally, in how engineering was talking with the business, how people were working together. Was there anything about this journey that was unexpected?

Tom: Not particularly. I think the biggest change that happened, which was possibly unexpected when we started, was the level of insourcing that we have undertaken to support the cloud-native operating models, the time to market, and the modern engineering teams. So, we used to be around 97 percent outsourced or something like that, in terms of building software that wasn't just vendor supplied. And for all that software now, we're more like 95 percent in-house. And so, that's quite a big change, and I think that probably surprised people that A) we needed to do it, and B) that we have done it, and relatively successfully got pretty wide-scale digital engineering functions across many markets now.

Emily: And why do you think that matters?

Tom: Because it gives us control of the roadmap, it gives us control of that time to market cadence, and it allows us to use the data that our teams understand and know about, and to share that with other markets. So, as I say, even though an engineering team might be in the UK, they can share what they've done, they can share the artifacts, they can share the data that's driven decisions and software activity with other markets within Vodafone. And that just improves that efficiency, again.

Emily: Do you think insourcing also improves customer satisfaction KPIs?

Tom: Certainly we've seen that. So, whether that's a correlation or causation is kind of for someone with more access to more data than I've got. But certainly, we've seen an increase in online sales, and our digital marketing is more data-driven. And that has happened in correlation with the in-sourcing of software engineering skillset, yeah.

Emily: Do you have any specific examples that come to mind in, maybe you are able to react in a way that wouldn't have been possible if you'd been using the old system?

Tom: I'm not aware of any specific examples, unfortunately.

Emily: Was there anything about the move to Kubernetes, to cloud-native, that you expected to be difficult, and wasn’t. So, that was easier than you expected?

Tom: That's a good question. I suspect the provision of multiple clusters. Kubernetes is difficult. It's a complex system, hence why there are so many cluster management vendor offerings available. And I think we chose a couple of partners early on in the journey to help us with that, and I think that really helped, and it made Kubernetes a little less scary for the software teams who were using it.

So certainly, I've heard feedback—this is anecdotal, rather than anything that's evidence-driven—actually, just being able to create clusters and deploy into them was easier than people had thought when they were learning about Kubernetes through the quick start tutorials and the like.

Emily: Was there anything that sticks out as being far more difficult than expected? The more unpleasant surprises?

Tom: I wouldn't necessarily call them unpleasant, but obviously there's going to be a transition period—which we're in—between the traditional data-center-centric networking and network security policies and concepts, and those that work with Cloud and cloud-native platforms like Kubernetes. And there have definitely been challenges in trying to apply the legacy approach to network security with a distributed and dynamic system like Kubernetes, where you can't give everything a static IP address or even have separate subnets within a cluster for segregation, for example. It has to be done in a different way. You can still apply the same controls, they just have to be done in a different way. So, I think that's one of a few challenges that we found that we've had to work through with different vendors, with engineering teams, and with our internal teams to try and update our guidance on how to apply those controls.

Emily: And to what extent have there been organizational challenges, and how have you gotten over those?

Tom: That's a tricky one to answer, really. I think it all comes down to the balance between understanding and buying into a strategy, but then applying that to application lifecycle and investment lifecycles. So, I think this is probably true for any company: just because a strategy says this is the thing to do, you got a roadmap for your portfolio of applications and services that you need to balance a limited budget. And so, that's been the biggest challenge, is to try and identify how much of each budget at various levels can be spent on strategic activity, and then for which services, and trying to keep that balance, and bearing in mind that there are lots of different things pulling on that same pot of money.

Emily: And what have you learned about managing that?

Tom: I think primarily that there needs to be a holistic view of strategic projects. It's quite difficult to put the onus on a local budget, to spend the money to do something strategic when the benefits are probably—and the business case is probably seen more widely than the individual budget area. But I think it differs between situations, and between markets, and what's happening. I think the primary thing is to understand the costs of the strategy upfront, and try and work those costs into whatever needs doing over the period.

Emily: A slightly different question, which is, is there anything you feel like in the cloud-native journey that you're still working on solving, that you haven't really figured out yet?

Tom: I'm not sure whether we haven't figured this out yet, but one of the things we're putting a lot of effort in at the moment, is the use of advanced data and analytics platforms to try and drive even more network automation, and network planning efficiencies. So, I think it was last year at Google Next, we announced a partnership with Google to make use of their data services. And there's a few projects ongoing within Vodafone to try and drive the amount of knowledge and useful information we can gather from the vast quantities of data we have about our services and the customers that use them because the more we can use that data, the more we can respond to customer need in a timely manner, whether that's reactively in terms of operational response or whether that's proactively in seeing trends that we can then meet a need that may be unsaid yet.

Emily: And if you were to talk to another engineering leader who was trying to push through the open door as you were saying, what advice would you give them?

Tom: The biggest bit of advice is to understand the current way of working for whichever area you're—is on the other side of the door, and understand their pain points because it's not always the same answer. So, generalizing, it may be that one area is more than happy to have a centralized global platform offering, whether that's within our data centers, or public cloud, or both. Another area, just the way it's managed, may require a more distributed model, where the services are offered on a more market specific level. And so, I think that that's the main thing, is to understand the specifics of that area that you're talking to because it will affect how you want to architect and onwardly deploy and manage that technology.

Emily: It would affect not just how you want to architect the technology, but also how you want to communicate what your plan is, right?

Tom: Absolutely. Yeah. So, in the first of the examples I gave, where an area might be happy with a centralized service, that probably means they're already using one. The way you would communicate that would be via that existing channel, if you like. Whereas on the flip side, that kind of channel may not exist, and therefore running the project or projects and communicating with stakeholders would be much more distributed.

Emily: At Vodafone was there ever any challenge selling it, not just over to the business side, but also selling internally inside engineering teams? Or was everyone pretty gung ho to do this?

Tom: No, there’s always challenges. I think again, it goes back to understanding the pain points of an area and understanding why things are the kind of as they are today, which I guess is general for things outside of technology and outside of Vodafone generally is. If you understand the position of the person you're debating with, then you're more likely to reach a common understanding than if you go into it with your own point of view and being unwilling to listen. So, I think that's the main thing is just being willing to listen, to understand pain points, and to be able to react to those within a strategy. You'd hope that it's flexible enough to be able to meet a wide range of needs without needing to necessarily change the overall vision.

Emily: How important do you think this cloud-native transition has been for Vodafone?

Tom: I think it's been crucial. I think we couldn't have done what we've done in the last five years without it. So, there's a video that our group CTO has posted on LinkedIn recently which highlighted a few things around improved mobile KPIs, we've got 4G in 21 markets, we've got the largest 5G in Europe, and all of those improvements from time to market I've already mentioned, we simply couldn't have done that without a modernization program to move to cloud-native across a number of our systems. So, yes, that's partly a technology thing, but also, it is such a cultural thing, and having that modern way of working where you have your modern engineering teams who are closer to the customer, but they're also—the different mindset of a modern engineering company where you're not afraid to try new things, and if you fail, you learn from them. And I think that's all part of what I would class as cloud-native, and that has been, like I say, it's been crucial for us to be able to get where we have been.

Emily: It's interesting to think cloud-native means if you fail, you learn from it. That's a fairly basic concept, and yet true. I can see how that is, sort of, part of being cloud-native.

Tom: Yeah, it's one of those things is quite a basic thing, but I think in traditional ways of working, the focus on the availability of systems and the performance of systems can blind everyone to the possibilities outside of that particular area of focus. And it puts pressure on people at all levels to try and minimize periods of downtime or periods of low performance. And over time, people become less and less willing to be able to try new things, through fear of failing because just the way people work it’s difficult to learn from those failings because it affects customers. And so, what cloud-native technologies enable because of the way things are orchestrated—things are dynamic, things are repeatable—it's very easy to try new things, and not affect all customers. Now, obviously, good software engineering practices help as well. But I think the cloud-native technologies and the ways of working really do support the whole “learn by failing” premise.

Emily: Do you think it would have been possible to get the customer satisfaction KPIs that you did, without moving to cloud-native, in any other way?

Tom: I think the only way you could have done is by a huge investment in people and the traditional technologies. It would have been a much more expensive and slower journey, in my opinion.

Emily: Anything else that you want to add about your experience moving to cloud-native?

Tom: No, I don't think so. I think one of the things—like I said before, the increase in automation, the increase in the modern technologies is just really helped with those customer affecting KPIs, and that has to be the drive for why you're doing it.

Emily: All right, just a couple more questions, then. What is your can't-live-without engineering tool?

Tom: Oh, that’s a good question. Probably Python. I think so many people use it either as a cross-platform scripting tool to be able to automate things and get on the first step towards cloud-native, or it's such a key part of many cloud-native tools like things like Ansible and other tools, and it's used hugely within our data analytics domain to try and drive the usefulness of the data. So, yeah, that's probably the one I’d choose.

Emily: And then this actually is the last question which is, how can listeners follow you or connect with you?

Tom: So, I'm on Twitter at @tomkivlin. I'm also on LinkedIn. So, I'm Tom Kivlin, working for Vodafone Group. I am a member of the telecom user group within the CNCF. So, you can find them on GitHub and also in the… I think it's the CNCF or the Kubernetes Slack. And yeah, happy to share experiences and keep learning.

Emily: Well, thank you so much. Again, this is Tom Kivlin, and we'll go ahead and wrap it up there. Thank you so much, Tom.

Announcer: Thank you for listening to The Business of Cloud Native podcast. Keep up with the latest on the podcast at thebusinessofcloudnative.com and subscribe on iTunes, Spotify, Google Podcasts, or wherever fine podcasts are distributed. We'll see you next time.

This has been HumblePod production. Stay humble.

View Details

This conversation covers:

  • Why many businesses are shifting away from analyzing total cloud spend (CapEX vs. OpEX) and are now forecasting spend based around usage patterns.
  • The difference between cloud-native, cloud computing, and operating in the cloud.
  • The delta that often exists between engineering teams and business stakeholders regarding costs. Travis also offers tips for aligning both parties earlier in the project lifecycle.
  • Common misconceptions that exist around cost management, for both engineers and business stakeholders. For example, Travis talks about how engineers often assume that business teams manage purely to dollars and cents, when they are often very open to extending budgets when it’s necessary.
  • Tips for predicting cloud spend, and why teams usually fall short in their projections.
  • Why conducting cloud cost management too early in a project can be detrimental.
  • Comparing the cost of the cloud to a private data center.
  • The growing reliance on multi-cloud among large enterprises. Travis also explains why it’s important to have the right processes in place, to identify cross-cloud saving opportunities.
  • How IT has transitioned from a business enabler to a business driver in recent years, and is now arguably the most important component for the average company.

Links:

  • Twitter: https://twitter.com/TravisWRehl
  • LinkedIn: https://www.linkedin.com/in/travis-rehl-tech/
  • Main Company Site: https://cloudcheckr.com
  • CloudCheckr All Stars: https://cloudchecker.com/allstars

Transcript

Announcer: Welcome to The Business of cloud-native podcast, where we explore how end users talk and think about the transition to Kubernetes and cloud-native architectures.

Emily: Welcome to the Business of cloud-native. I'm your host, Emily Omier, and I'm here today with Travis Rehl, who is the director of product at CloudCheckr. Travis, I just wanted to start out, first of all, by saying thank you for joining me on the show. And second of all, if you could just start off by introducing yourself. What you do, and by that I mean, what does an actual day look like? And some of your background?

Travis: Yeah. Well, thanks for having me. So yeah, I'm Travis Rehl, director of product here at CloudCheckr. What that really means is, I have the fun job of figuring out what should the business do next in relation to our product offering here at the business. That means roadmap, looking at the market, what are customers doing differently now, or planning to do differently over the next year, two years or so, on cloud? What their cost strategies are, what their invoicing and chargeback strategies are, all that type of fun stuff, and how we can help accommodate those particular strategies using our product offering.

Sort of, day to day, though, I would say that a bunch of my time during the day is spent talking to customers, figuring out where they are in their cloud journey, if you will, what programs or projects they may have in flight that are interesting, or complicated, or they need help on. Especially making any sort of analysis help in particular, and then lastly, taking all that information and packaging it up neatly, so that the business can make a decision to add functionality to our product in some way that can assist them move forward.

Emily: The first question I wanted to ask is actually if you could talk just a little bit about the distinction between cloud-native, and cloud computing, and operating in the cloud. What do all of those things actually mean, and where's the delta between them?

Travis: Sure. Yeah so, it's actually kind of interesting, and you'll hear it a little bit differently from different people. In my background, in particular—I used to run an engineering department for a managed service provider. And so we used to do a lot of project planning of companies as to what's their strategy for their software deployment of some kind on cloud. And typically the two you see for, say, cloud-native versus operating in the cloud, operating on the cloud is very atypical.

You'd associate that to something like lift and shift—probably hear about a lot—the concept of taking your on-prem workload and simply cloning it, or taking it in some way and copying in some way, on to the cloud-native vendor in particular. So, literally just standing up servers of clones of hard drives and so forth, and emulating what you had on-prem, but on the cloud. That's a great technique for moving quickly to cloud. That's not a great technique if you want to be cloud-native. So, that's really the big segue for cloud-native, in particular, is you want to build a software solution that takes advantage of cloud-only technology, meaning serverless compute resources, meaning auto-scaling different types of services themselves, stuff you probably didn't have when you're on-prem originally, that you now have, you can take advantage of on the cloud. That's almost like a redesign, or reimplementation around those models that cloud itself provides to you. That, to me, is the big difference. And I see oftentimes that gap-wise, many companies who are starting on-prem, they will do the migration to cloud first, the lift and shift model, and then they will decide, “Hey, I want to redesign pieces of it to make it more cloud-native.” And then you'll see startups who don't have on-prem at all, they will just go into cloud-native from the get-go.

Emily: Of course, CloudCheckr specializes in helping with costs among some other things, but how do costs fit into this journey, and what sort of cost-related concerns do companies have as they're on this cloud journey?

Travis: Yeah, so there's a few. I would actually say that years ago—just to clarify, the discussion has changed over the last few years—but years ago, it started with CapEx versus OpEx costs, specifically for purchasing of your IT services. On-prem, you'd probably purchase up-front a bulk number of VMs or servers or otherwise, for a number of years, and so be a CapEx cost. When you moved over to cloud and more of this, usage-based, model kind of threw a lot of people for a loop when it came to more OpEx usage space models. AWS, Azure, GCP have helped in that regard with things like reserved instances for companies who are more CapEx oriented as well, but in terms of the initial years ago, a big hurdle was communicating that difference and how the business may pay for these services. And a lot of people were very interested in moving to OpEx back then, in particular.

When it came to how do you take into account all these cost-related changes the business may go through, one of the big ones that I see most recently is around the transference and storage of data. In the past, it would have been about how much money total am I going to spend on the cloud itself. Now, it's about what am I forecasting to spend based off of those usage patterns. It's a bit easier to forecast those things when you have servers that run for a period of time, but when you have usage patterns for data ingestion, for data transfer, for servers spinning up and spinning down and scaling out horizontally, that pattern becomes a bit more fluid. So, that's typically a conversation that comes up quite often. That's the type of thing that CloudCheckr and products like us can help with.

Emily: Do you think there's any delta between how engineering teams think about and understand cost, and business stakeholders think about and understand costs?

Travis: I would say yeah, there's a fairly significant difference there. I would say that engineers initially care a little bit less about cost. They have an objective. They're trying to solve for a project or goal they're trying to achieve. And so as a result, they're saying, “How can I achieve that quickly?” And that changes over time as the product becomes closer to production-ready. They may say, “Okay, now I want to optimize a bit more what I built.” Whereas the business is more thinking about, “I have this project plan in front of me, how do I go as quickly as possible, without incurring so much costs that it goes over budget, or I don't foresee a particular budgetary circumstance occurring?” It's slightly different mentalities, to be quite honest with you. What I see the most of is that they do align towards the end of the project, or the steady-state of the project, when the team has delivered the thing, whatever that may be, but need to make a decision quite quickly as to how do they want to either cap those costs moving forward or come up with an appropriate sort of budgetary model that scales linearly or otherwise, that both sides can agree upon?

Emily: And are there any best practices that you can think of for how that gap can be bridged earlier in the process?

Travis: Yeah. What’s interesting, you're seeing, probably in the last couple years now—as you've probably heard, “Cloud Centers Of Excellence.” I'm assuming you've heard that term?

Emily: Mm-hm. Yep.

Emily: Yeah. So, one of the big amplifiers of that is bringing in your financial team to that planning model for how you want to deploy and manage cloud resources, so they have a say early on in the process, but also, it's more about communicating what the plan needs to be. Typically, I see the engineering team—the product manager, right—going off and saying, “We need a build and go fast.” And then it's the financials—accounting or otherwise—that says, “Hey, what happened to this project once you delivered it?” To bring them in earlier to that conversation via a CCOE sort of methodology—Cloud Center Of Excellence methodology—I see helps the most. It’s really about communicating first.

The second, though, is about setting corporate guidelines for what people are allowed or not allowed to do. Shadow IT has been a problem for a very long time for businesses. Cloud actually, in my opinion, can amplify that problem because of how easy it can be to start spinning up resources, and doing projects, or otherwise. So, it's pretty important for central IT, or for that Center Of Excellence group to set the standards of how the business needs to operate from that point forward, but then allowing different organizations, business units, or groups to deploy and manage at the cadence they need.

Emily: What do you think are some misconceptions on both sides about cost management? So, misconceptions from engineers and misconceptions from business stakeholders?

Travis: I think, as a product person, I have the fun job of being someone in the middle of it all. I think engineers have a misconception that the business side is managing to dollars and cents. That can often be the case. So, if you're an engineer, and you are a creative type, and you want to build something, and you have someone out there saying, “No, no, no, you can't do it because of XYZ thing,” cost or [unintelligible], or otherwise it can create this barrier between the two of them. But what I have learned, though, is that from an engineer communicating to a business person, especially on this subject, documenting and communicating the impact, the positive impact, and sometimes negative impact, but positive impact I'd recommend, on the business for why the project will enable your teams, or get you into a market faster, or solve a corporate problem that's been around for a while.

The business side is very open, I found over the years, to hearing out those reasons, and agreeing to extending costs, or being more lenient on the cost model of a project so it can go faster. Likewise, though, on the business side, when they're working with engineers, it can be a little bit fraught when they think of, say if you've got a project moving and a business team says, “You have a budget for this and a line item,” and they say, “What's the project plan? When are you going to deliver? What's the velocity for delivering?” Cloud lets engineers do a bunch of new things they probably haven't had in the past, like faster deployment times to the product or a project in question; being able to react, or be proactive, however way you want to manage to it, for a customer requests, or otherwise. And as a result, it can be easy for that business person to expect a very linear approach to product development or delivery of services. And so for them, it's more about seeing and working with engineers to see how, “Hey, maybe the business can be improved by doing things a little bit differently than they had in the past.”

Emily: Yeah. So, basically, what you're saying is, as long as the business is understanding this expense as an investment in something, usually it's not a big problem.

Travis: Yeah. But it really comes down to communicating that. It can be so hard for an engineer, to say, “Hey, if I deliver this new pipeline, then I can deliver code, converge code, deploy code, test code faster.” That may not translate really well to a business person, and that's kind of the, I would hope, the role of product within your organization, or their organization, to help do that translation because if you're able to deploy quicker, that means you can solve customer problems faster, your time to market’s faster, you can make changes to your cost model faster, with very little risk. But that's the communication gap that typically resides.

Emily: And what do you think are some surprises that come up on the cloud journey regarding—we've already talked about how this is a shift from CapEx to OpEx, but what do people not fully understand about what those OpEx costs are going to be for example, or how it's going to work at the beginning of the journey versus at the end?

Travis: Yeah, so at the beginning, I think there's an interpretation that—let's say you're spinning up a project for a very first time, and you're typically used to that CapEx model, and the team delivers their budget to you, the business person, and they’re like, “It’s going to cost us, round numbers, like$1,000 this month to get started.” And then by month two, they've scaled up, and they're two grand or three grand or something, whatever that number needs to be. The business is typically not prepared or is ready to be prepared to understand that it will change; the costs will change. It never is exactly what you ever thought it was going to be at to start. You can estimate all you want, there's tons of different calculators out there that helps you get close enough, but it's never going to be the number you always envisioned it to be. And that's just something you have to live with and get used to. It's more about the later on in the project that matters more.

So, I think up front, it's very useful for project leaders to also communicate to the business that this is an estimate. We think it's going to start here, but we expect it to grow. And then, have good caps for yourself, almost like milestones to say, if you're getting closer to a particular cap, to a particular budget, maybe you should sit down for a moment and say, “Hey, are we doing this the wrong way, or can we do it a little bit differently?” Later on in the project, though, it gets kind of interesting. So, in my past life, I worked at a company where we did a lot of eCommerce implementations for other vendors, for other businesses, and one of the things always surprised people the most was data transfer costs.

So, when we will be scoping a project, we would say, “Okay, here's the traffic we would expect to hit your new eCommerce solution, here's the type of things we expect them to do, here is the type of content, they need to load from the system, and thus, here's how much you're going to spend in transfer rates from the servers or otherwise—or CDN or otherwise—to the end-user.” And then what always hit us every single time is, like, Black Friday, or some kind of promotional event a marketing team would do, where we see a ton more traffic coming onto the system. And suddenly, all those atypical usage statistics would spike, you would be sending more images out to users, you'd be ingesting more traffic than you're anticipating, and you'd have more storage of their user information or otherwise, on your back end. And so, that sort of scaling model of usage patterns becomes more common over time because you see it more often. I like to put sort of, like, buffers on those particular areas of the budget, that say, “Hey, here's how much I think I'm going to spend on EC2, or computes, or VMs, or what have you, but I know that we ingest a lot of images from end-users, and we have to take in those images all the time, and I know we're going to run four promotions in a given year, and I should expect a 40 percent increase on those time periods. So, maybe my OpEx costs should be increased by a certain budget to accommodate that.” That type of planning has to start, has to become normal to happen between the project team and the accounting group, or financial group of some kind. So, there's some more of a clear agreement as to what realistically the budget needs to be long term.

Emily: Do actual cloud costs ever come in less than estimated?

Travis: I would say it's more often over what's estimated. The only times I see it coming in less is when the team delivering the solution implements a really sophisticated sort of data transfer strategy, or storage and archival strategy than expected. Maybe usage is less than you expect as well. Or, everybody should always be planning for some buffer, so if you've got a project that you think it's going to cost you $800 a month, maybe you should tell the business it’s $1000. [laughs]. Sort of overachieve where you think it’s appropriate.

In that case, you're always going to come in lower than you expect. I actually think that cloud costs will always increase. That's just natural to do so. As you do more on cloud, you store more, you transfer more, you need more compute; it will always increase. I think it's more important to analyze the rates in which you are increasing. If you are spiking and it continues to spike, then there is a problem. If you're steadily growing, but your user account is growing, your revenue is increasing as well, and it's linear in some fashion, in that regard, that's a good thing. That doesn't mean something's wrong. But if you're spiking and things are consistently spiking, then you have a problem.

Emily: Just like any other part of your business. If you're hiring more people, hopefully, it's because it means your business is growing. But you're also spending more—

Travis: Exactly.

Emily:—obviously. You were also saying sometimes the costs end up being less than expected because the company has invested in cost management strategies. Why—

Travis: Yeah—

Emily: Oh, go ahead.

Travis: I was just going to say, I would say that companies don't typically start there. At least in the past, they haven't started first food forward, where they're saying, “I want to keep costs in check first, prior to doing the project.” Typically, they do a project, then they find that oh, they did too much, and now we need help. That's the normal I see.

Emily: Yeah, in fact, I was going to ask, are there disadvantages to doing cloud cost management?

Travis: In general?

Emily: Yes, or scenarios where you would say, “Maybe this isn't what you're supposed to be focused on?”

Travis: Yeah, I think it depends on the stage of either the entire company or the individual business unit or project. And what I mean by that is, if you're Google, you're probably managing a very small team, like they’re a startup, or your startup yourself. And so, in the very early onset of any product delivery, cloud or not, it's typically more about volume and usage patterns of your end-users on the system. And you will spend money to get there, you will put in long hours to make it happen. Don't create barriers yourself, to achieve the goal that you have in mind for your team, your project, or otherwise.

But then as the project matures, you typically see people sort of step back and say, “Okay, we made a great lot of great strides. We've hit a lot of our milestones, we have the user base we need, how do we tailor back a bit so that things become a bit more normalized and consistent on the delivery?” And so early on, I don't see a lot of people putting a lot of effort into cost savings, cost optimizations, and recommendations, just because things are fluid, dynamic, chaotic, whatever word choice you want to use for yourself at the time. But at some point in any business, you will come to the conclusion of, “Okay, we've done a lot of things; a lot of stuff. How do you then efficiently optimize what you have already delivered?” And that's where tools like CloudCheckr and others come into play to help you figure out that cost optimization and that strategy.

Emily: Yeah. And can you think of any examples where you think, maybe this company would have paid less if they had done cost management, but maybe they wouldn't have accomplished x?

Travis: Yeah. Without using a customer name, but to be honest, I forget their name offhand; it’s about a year ago, we had a particular scenario where they're moving from on-prem. They were a video provider, and they were moving from on-prem to cloud. And in the move they had, it was something like $400,000 per month in wasted resources. Meaning they had IP addresses not attached to particular servers, or load balancers, they had storage volumes that are no longer in use, not even being backed up or anything, just kind of live off to the side, that people kind of forgot about. It was like $400,000 a month in that stuff. It was a lot of money.

And they actually made a decision in the middle of the project, to back up and slow down for a moment and clean up all those things, and it probably cost them a couple months or so—because of the size and volume they're driving at—it probably cost them a couple months of their internal resource time, or otherwise to remediate those loss cost savings, and then continue the project forward with better guardrails. If I remember correctly, they actually missed a significant event day. Like, it wasn't Black Friday, it was some media day they had at the company they wanted to leverage a new system for so that it had a more performant experience with their customers to see videos and content they were delivering. And they had to leverage their older system at the time, which actually cost them more money at the moment because they had to spin up additional resources on their old system to handle that load, and it just wasn't optimal experience for themselves to do that. I think they broke even at the end of the day. I would have just gone faster on the new thing in that scenario, but that's how it played out for them.

Emily: So, basically, what you're saying is that it's not always a really simple decision. There's pluses and minuses, and everybody has a limited amount of time, resources, and focus, so you have to decide what's most important at this moment for your organization.

Travis: That particular daily experience is near and dear to my heart because that's the majority of my job. There are a lot of opportunities in front of you, a lot of bad things you could do, too, sometimes you're not going to get everybody happy, either. You need someone who's going to make a decision and move quickly, especially if you're a business that's on cloud in a competitive market that's trying to grow. You don't have months or years—sometimes even weeks—to make that type of decision. You kind of have to have all the data in front of you and say, “Here's the best one.”

Cloud makes that harder to be quite honest with you. The speed at which you can deliver functionality, and your competitors themselves, the speed at which they can change means that you need to be very confident in what you can do and what your team can do, as well as understand what the risks are in front of you so you can make a decision quickly. If you're not ready to do that, then there may be problems ahead for you because as the market continues to go in that direction, that's just going to be a soft skill required moving forward for a lot of people.

Emily: Do you think that the cloud is less expensive than data center? Do you think that cloud-native is less expensive than say someone doing a lift and shift?

Travis: Yeah, so it depends a little bit. So, I can start with cloud-native and then work my way back to private cloud. It will be more expensive upfront, to go cloud-native typically, because you're building it from scratch, or redesigning a significant portion most likely of an existing application stack. So, the upfront overhead to go cloud-native is always going to be higher, especially when you factor in things like labor costs. However, though, once you have made that transition, costs can be tremendously lower. So, it's really more so about, is the business willing to take on the effort to make that leap upfront, or is there only so much that’s palatable [laughs] at the time to move forward.

But typically, it's a bigger spike upfront with much cheaper experience late game. When it comes to things like on-prem, or private cloud, or even hybrid cloud to some extent, I actually think if you have workloads that are very consistent, if you know exactly how much storage you're going to need, how much compute power you need, it's very stable, very consistent, it's probably cheaper to go to your private data center because you can purchase in bulk at the best rates for the time period you need for that specific thing, because you know exactly what it's going to be. That is only the case for a small subset of businesses in the world, who even have the analytics to come up with that data. But if you do, it's completely viable. Was it Dropbox, I think, recently—not recently. A couple years ago now, actually—they went off of their cloud provider. I think they're on AWS at a time, offhand. And they went to their own private data center. That could be butchered in this, slightly, but you get the point, is they knew exactly how much storage they need. And they found it to be cheaper to do it themselves on their own private data center. So, they did. Not a lot of companies can do that. But it's completely an option out there.

Emily: What do you think the stakes are? And by which I mean, how big a piece of the pie is IT budgets?

Travis: So, what funny is five, six years ago, if you were to ask me that question—this is back when even I worked more true to form IT—IT was not a business driver. It was overhead. It was, make sure my email is still on, and I have computers, and I can do these things, and connect the networks together. And it wasn't driving business decisions, type of a thing. That has radically changed over the last few years. I would actually say that IT is probably the most important aspect of your business now.

Without a really strong IT department, you can't move fast on product development. Without a strong IT department, your engineering teams are—the velocity of the deliverables is slower. You can't leverage all these great new stuff that's coming out. So, IT now, to me, is probably the most important thing a business can truly manage appropriately. When it comes to the total budget, it has greatly increased because more responsibilities have also been placed upon IT.

They now have to manage to different usage patterns from forecast, different budgets. They have to manage to new skill sets requirements for cloud in particular, and other different types of niche technologies that your business may need. That means that IT themselves has to either staff up or purchase technology solutions like CloudCheckr to accommodate that growth in the business requirement. So, the IT budget is greatly increasing, but for good reasons. It's because the business demands nimbleness, and only a strong IT department can deliver that.

Emily: Does that mean though, that it's more important than ever to figure out where you're wasting resources, or at least be strategic about where you're spending those IT resources?

Travis: I one hundred percent believe in the strategy of identifying and managing to cloud waste should be in the forefront of pretty much every conversation, among other items. When it comes to focusing entirely on cloud waste, I don't think that's the right decision. I think there's a way that manage to that, that the business will accept while allowing for the right velocity of improvements to products or services, but when it comes to defining—really this comes down to me is the Center Of Excellence, at the end of the day. The Center Of Excellence should be that IT is either the owner of, or is a prominent member of. And they should be defining what's the atypical cost-saving strategy? What is the security profiling? What is the particular way you want to tag your resources, or otherwise, so that you can allocate them to the right business unit, or project code, or etcetera, for that cost analysis? That is the objective of the Center Of Excellence is defining the strategy. Cost savings is a piece of—a significant piece of—but a piece of the overall strategy that you should have an answer to, but it should not be the primary driver.

Emily: Anything else that you'd like to add?

Travis: In relation to Centers Of Excellence?

Emily: Just in relation, actually, to this whole discussion about cloud-native, costs, etcetera?

Travis: Yeah. Really, I want to speak to more on the large enterprise side, and/or smaller companies who are maybe being acquired, and having to manage themselves inside of a larger organization all of a sudden. It's becoming fairly normal to see two cloud providers inside of a larger organization. And the normalcy that I've seen is acquisitions: a larger business purchasing a smaller one, maybe the larger ones on AWS, but a smaller one has chosen GCP. And no company is going to tell the new acquisition to refactor themselves and shift over to Amazon. That's a—millions of dollars will be spent [laughs], and lots of time spent to accommodate that request, so no one's really going to answer it.

Instead, what they will, though, is central IT will be asked to manage both. So, multi-cloud is becoming a thing. It's becoming more and more important to have the same business strategy applicable to multiple cloud vendors at once. It's important to have the right tooling in place so that as you onboard different cloud vendors than you're normally used to, at least the terminology stays the same, the functionality stays the same because you have the right tools that can help make that translation for you. Without that right process in place, it will mean that your organization will have to incur overhead to help manage always to different costs, different ways to optimize your costs, or otherwise. And so multi-cloud tooling is becoming more and more important for larger orgs, especially as your atypical business processes pan out.

Emily: How exactly, actually, does multi-cloud impact to this sort of cost equation?

Travis: Yeah, so there's two trains of thought when it comes to how multi-cloud impacts a project and then how the project [eventual] cost savings. So, there's the one that I personally believe does not happen that often, but it's worth mentioning, where you have a project that's deployed—you know, software deployed, application stack deployed—and that's leveraging a small piece of it from a different cloud vendor. So, maybe 90 percent of the stack is deployed on AWS but uses Google Cloud for some analysis engine they may have, or some particular tool. So, now you've got to combine these two things together. That happens, but rarely.

The alternative to that is you have discrete projects who have their own cloud vendors that they use 100 percent of but you now need to manage the business unit, owns both those projects. And how do you then combine those two things, but there is no project dependency between the two? In both cases, there needs to be a strategy employed for identifying the individual resources. So, in AWS or otherwise, you have a tagging strategy based off of project code, or business unit, or reporting lines, or however the business needs to manage to it.

The second is, you need to come up with and identify consistent ways you want to save money. Every cloud vendor may have a slightly different functionality because they need to differentiate among themselves; that's normal. But everybody should be turning off servers from 7 p.m. at night until 6 a.m. the next day, when the workforce is typically not working—the internal systems, pre-production type systems—if you can. And that's true to form, regardless if you’re Amazon, or Azure, or GCP. And so you need tooling and the right process in place to identify those cross-cloud cost-saving strategy options available to you, and need to normalize the way you want to implement on it. If you don't do that, what typically will end up happening is you're going to be working out of multiple consoles with different terminology, and different settings you're not used to, and different ways to implement it, and somewhere something will fall through the cracks, and you will not have consistency in your cost-saving strategy.

Emily: All right, one more question. What's your favorite can't live without engineering tool?

Travis: Oh, I can't say that. Because I'm going to get yelled at. [laughs]. If I had to take off my CloudCheckr hat and put my personal hat on for a brief moment, I would say log analysis tools like Sumo Logic, like Splunk, is my favorite thing, period. Back when I was more in engineering side of the house, before I was in product, I would spend hours looking through log analytics. There's just so much you can do with those tools nowadays to identify, come up with really cool analysis products.

That's actually one of the guiding principles I see for CloudCheckr is people really want to have fun tools they can play with, and see, and get their hands on, and come up with new conclusions to answers they may not have had in the past. That's what I see those log analytics products doing. CloudCheckr is following a very similar route, and that's the ethos that I'm instilling on our product team is our users, our customers, they should have a fun product to be able to get into and play with their data, and use in different ways to come to new conclusions they’ve never seen before, especially on cost analysis. So, those are my favorite products and how we integrate them into our product decisions.

Emily: Where can listeners connect with you?

Travis: Yeah, so I'm on LinkedIn. Just Google my name Travis Rehl, CloudCheckr, feel free to contact me; happy to chat. CloudCheckr, if you're a CloudChecker customer, we have an All-Stars program at cloudchecker.com/allstars, which enables you to have a Slack workspace with us, which myself and my team run so you're able to chat, have conversations like this, or otherwise. I also post on Twitter, but be honest with you not that often. So, if you're looking to reach out, I’d say LinkedIn is probably your best bet.

Emily: Thank you again. This is Travis Rehl. And thank you for joining us on the Business of cloud-native.

Travis: Sure. Thanks for having me.

Announcer: Thank you for listening to The Business of cloud-native podcast. Keep up with the latest on the podcast at thebusinessofcloudnative.com and subscribe on iTunes, Spotify, Google Podcasts, or wherever fine podcasts are distributed. We'll see you next time.

This has been HumblePod production. Stay humble.

View Details

Some of the highlights of the show include:

  • The difference between cloud computing and cloud native.
  • Why operations teams often struggle to keep up with development teams, and the problems that this creates for businesses.
  • How Dave works with operations teams and trains them how to approach cloud native so they can keep up with developers, instead of being a drag on the organization.
  • Dave’s philosophy on introducing processes, and why he prefers to use as few as possible for as long as possible and implement them only when problems arise.
  • Why executives should strive to keep developers happy, productive, and empowered.
  • Why operations teams need to stop thinking about themselves as people who merely complete ticket requests, and start viewing themselves as key enablers who help the organization move faster.
  • Viewing wait time as waste.
  • The importance of aligning operations and development teams, and having them work towards the same goal. This also requires using the same reporting structure.

Links:

  • Company site: https://www.mangoteque.com/
  • LinkedIn: https://www.linkedin.com/in/dmangot/
  • Twitter: https://twitter.com/DaveMangot
  • CIO Author page: https://www.cio.com/author/Dave-Mangot/

Transcript

Announcer: Welcome to The Business of Cloud Native podcast, where we explore how end users talk and think about the transition to Kubernetes and cloud-native architectures.

Emily: Welcome to The Business of Cloud Native. I'm your host, Emily Omier, and today I am chatting with Dave Mangot. And Dave is a consultant who works with companies on improving their web operations. He has experience working with a variety of companies making the transition to cloud-native and in various stages of their cloud computing journey. So, Dave, my first question is, can you go into detail about, sort of, the nitty-gritty of what you do?

Dave: Sure. I've spent my whole technical professional career mostly in Silicon Valley, after moving out to California from Maryland. And really, I got early into web operations working in Unix systems administration as a sysadmin, and then we all changed the names of all those things over the years from sysadmin to Technical Infrastructure Engineer, and then Site Reliability Engineer, and all the other fun stuff. But I've been involved in the DevOps movement, kind of, since the beginning, and I've been involved in cloud computing, kind of, since the beginning.

And so I'm lucky enough in my day job to be able to work with companies on their, like you said, transitions into Cloud, but really I'm helping companies, at least for their cloud stuff, think about what does cloud computing even mean? What does it mean to operate in a cloud computing manner? It's one thing to say, “We're going to move all of our stuff from the data center into Cloud,” but most people you'll hear talk about lift and shift; does that really the best way? And obviously, it's not. I think most of the studies will prove that and things like the State of DevOps report, and those other things, but really love working with companies on, like, what is so unique about the Cloud, and what advantages does that give, and how do we think about these problems in order to be able to take the best advantage that we can?

Emily: Dive into a little bit more. What is the difference between cloud computing and cloud-native? And where does some confusion sometimes seep in there?

Dave: I think cloud-native is just really talking about the fact that something was designed specifically for running in a cloud computing environment. To me, I don't really get hung up on those differences because, ultimately, I don't think they matter all that much. You can take memcached, which was designed to run in the data center, and you can buy that as a service on AWS. So, does that mean because it wasn't designed for the Cloud from the beginning, that it's not going to work? No, you're buying that as a service from AWS.

I think cloud-native is really referring to these tools that were designed with that as a first-class citizen. And there's times where that really matters. I remember, we did an analysis of the configuration management tools years back, and what would work best on AWS and things like that, and it was pretty obvious that some of those tools were not designed for the Cloud. They were not cloud-native. They really had this distinct feel that their cloud capabilities were bolted on much later, and it was clunky, and it was hard to work with, whereas some of the other tools, really felt like that was a very natural fit, like that was the way that they had been created. But ultimately, I think the differences aren't all that great, it just, really, matters how you're going to take advantage of those tools.

Emily: And with the companies that you work with, what is the problem or problems that they are usually facing that lead them to hire you?

Dave: Generally the question, or the statement, I guess, that I get from the CIOs and CTOs, and CEOs is, “My production web operations team can't keep up with my development teams.” And there's a lot of reasons why those kinds of things can happen, but with the dawn of all these cloud-native type things, which is pretty cool, like containers, and all this other stuff, and CI/CD is a big popular thing now, and all kinds of other stuff. What happens, tends to be is the developers are really able to take advantage of these things, and consume them, and use them because look at AWS. AWS is API, API, API. Make an API call for this, make an API call for that.

And for developers, they're really comfortable in that environment. Making an API call is kind of a no brainer. And then a lot of the operations teams are struggling because that's not normal for them. Maybe they were used to clicking around in a VMware console, and now that's not a thing because everything's API, API, API. And so what happens is the development teams start to rocket ahead of the operations teams, and the operations teams are running around struggling to keep up because they're kind of in a brand new world that the developers are dragging them into, and they have to figure out how they're going to swim in that world.

And so I tend to work with operations teams to help them get to a point where they're way more comfortable, and they're thinking about the problems differently, and they're really enabling development to go as quickly as development wants to go. Which, you know, that's going to be pretty fast, especially when you're working with cloud-native stuff. But I mean, kind of to the point earlier, we built—at one of the companies I worked at years ago—what I would say, like, a cloud environment in a data center, where everything was API first, and you didn't have to run around, and click in consoles, and try to find information, and manually specify things, and stuff like that; it just worked. Just like if you make a call for VM in AWS, an EC2 instance. And so, really, it's much more about the way that we look at the problems, then it is about where this thing happens to be located because obviously cloud-native is going to be Azure, it's going to be GCP, it's going to be all those things. There's not one way to do it specifically.

Emily: What's the business pain that happens if the operations team can't keep up with the developers? What happens? Why is that bad?

Dave: That's a great question. It really comes down to this idea of an impedance mismatch. If the operations teams can't keep up with the development teams, then the operations teams become a drag on the business. There's so much—if you read the state of DevOps reports that are put out by DORA research and—I guess, Google now, now that they bought it—but they show that these organizations that are able to go quickly: the organizations that are able to do deploys-on-demand, the organizations that are able to remediate outages faster, all those things play into your business's success.

So, the businesses that can do that have higher market capitalization, they have happier employees, they have all kinds of fantastic business outcomes that come from those abilities, and so you don't want your operations team to be a drag on your organization because that speed of business, that ability to do things a lot more easily, let's even call it like a lot more cloud-native if you want, that has real market effects. That has real business performance impacts. And so, if you look at the DevOps way of looking at this—like I said, I've been pretty involved in the DevOps movement—really the DevOps is about all the different parts of the organization working together in concert to be able to make the organization a success. And the first way of DevOps, you're talking about systems thinking, you're looking at the overall flow of work through the system, and you want to optimize that because the faster we can get work flowing through the system, the faster we can deliver new features to our customers, bug fixes to our customers, all the things that our customers want, all the things that our customers love. And so if you're going to optimize flow of work through the system, you definitely don't want work slowing down inside the operations part of the system. That's bad for the business, and that's bad for your business outcomes.

Emily: And how do you think companies realize that this is a problem? I mean, is it obvious or not?

Dave: I think it's one of those things like I always talk to people about process right? When do we want to introduce process? a lot of startups are like, “We need more process here, we need more process there.” And my advice to everybody is always, use as little process as possible for as long as possible, and when you need that process, it will make itself known. The pain will be so obvious that you'll be like, “Okay, we can't do this anymore this way. We've run this all the way to the end, and now we have to change things, and now we have to introduce process here.”

And I think that it becomes pretty obvious to certainly the companies that I work with. At one point where they're like, “This isn't working, I'm getting my development leadership coming to be saying ‘I’m waiting for this, I'm waiting for that. I'm waiting for this. I'm waiting for that. I don't have permission to do this. We're being blocked here.’” all the things that you don't want to be hearing from your development leaders because what they're expressing is their pain of being inhibited; their pain of being slowed down. And I think it's just, like, with the process thing, I think at some point, the pain becomes obvious enough that people say, “We have to do something.”

I remember talking to one company, and I was like, “Well, what do you want out of this engagement? What's your end goal?” And they said, “We'd like it for our developers to show up every day and be really happy with the environment that they're working in.” And so, you can hear it there, right? Their developers are not happy. People are coming from other companies, and they're going to this company—and I certainly won't name who they are—but they're going to this company, and they're saying, “Hey, when I worked at this other place, I didn't have all this. I didn't have all these things stopping me. I didn't have all these things inhibiting me.” And so that's why, what I said in the beginning, it's things that I'm hearing from CEOs, and CTOs, and people in those positions because at some point that stuff is just bubbling up and bubbling up, and the amount of frustration just really makes itself known.

Emily: Why do you think COOs care how happy their developers are?

Dave: Well, I mean, there's tons of studies that show the happier your developers are, the more productive they are. I mean, look at the Google rework stuff about psychological safety. Google discovered after hiring a professional psychology researcher to determine who were their highest performing teams, their highest performing teams weren't the teams that had the most talented engineers; it wasn't the people who went to MIT; their highest performing teams weren't the ones who had the best boss, or the coolest scrum master, or anything like that. Their highest performing teams were the teams that had the most psychological safety. People who were able to operate in an environment where they felt free to talk about the things that maybe weren't going well, things that could be improved, crazy ideas they had to make improvements, stuff like that.

And I don't think that you can be on a team that's unhappy and feel like there's a lot of psychological safety there. And so, I think those things are highly correlated to one another. So, I mean, obviously, the environment that's necessary for psychological safety goes far beyond whether or not my Kubernetes cluster is automatically deploying my Docker containers; that's certainly not the case. But I think it's important to recognize that if developers are in an environment where they feel empowered, and they're not being inhibited, and they can really focus on their work, and improving things, and making things better, that they're going to produce better work, and that's going to be better for companies and certainly their business outcomes.

Emily: And bringing it back to cloud-native a little bit. Can you connect for me how a cloud-native type architecture helps bring operation teams up to speed or helps remove these roadblocks?

Dave: Yeah. Well, I think it's a little bit of the reverse, right? I think that the successful operations teams are the ones who are enabling these cloud-native ways of looking at the world. I think there used to be this notion of, if you want something from operations, you open up a ticket, and then operations goes and they do the ticket, and then they come back to you and say, “It's done.” And then, this never-ending cycle of sending off something and then waiting, and then sending off something, and waiting.

And in cloud-native environments, we don't have that. In cloud-native environments, people are empowered and enabled, to go off and deploy things, and test things, and remediate things, and do dark launching, and have feature flags, and all these other things that, even though we're moving quickly, we can do that safely. And I think that's part of the mind shift that has to happen for these operations teams, is they need to stop thinking about themselves as people who get things done, and they need to start thinking about themselves as people who are enabling the whole organization to go faster, easier, better. I always talked to my SREs—Site Reliability Engineers—who used to work for me, and I'd say, “You have two responsibilities and that's it. And this is, in order, your first responsibility is to keep the site up.” That sounds pretty normal, right? That's what most operations teams feel like they've been tasked with. And I'm like, “Your second responsibility is to keep the developers moving as fast as possible.”

And so really, when you start taking that to heart, keeping developers moving as fast as possible, that's not closing tickets as fast as you can. That's not keeping the developer moving as fast as possible, that's enabling developers to have self-service tools, and have things where they want to get something done and it's very painless for them to do that. We used to launch EC2 instances at one of the companies I was working with, where we had gotten it down to a point where you just said what kind of machine you wanted, and then that was it; and you were done. And everything else got taken care of for you: all the DNS, all the security groups, routing, networking, DNS, like, everything was all taken care of, all the software was loaded. There wasn't anything to do but say what you wanted.

And we actually were able to turn that tool over to the developers so they could launch all their own stuff. They didn't need us anymore. And I think that's really creating these cloud-native ideas. Certainly, a lot of that stuff is part of the cloud-native tooling, now. This was a few years ago, but really it's enabling the developers to go as fast as possible. We could have said, “Hey, you want a machine? Open up a ticket, and it's so easy for us to spin up a machine.” But we didn't do that. We took it to the next level, and we empowered them, and we allowed them to go quickly. And that's really the sort of mental shift that the operations teams have to make. How do we do that?

Emily: I have to say, I have never been a developer, but whenever anyone talks about this process of submitting a ticket and waiting for it to get addressed, it just sounds like hell.

Dave: Yeah. Well, if you look at it from a lean manufacturing, Toyota kind of thing, all that wait time is waste. In lean, they call that waste. It's a handoff: there's no work being accomplished during that time, and so it's waste in the system. And so Toyota is always trying to move towards—I can't remember they call. It, I think it was, like, one piece flow, or something like that where basically you want work to be happening at all times in the system, and you certainly don't want things sitting around.

And so, developers don't want that either. Developers want to put things out there. They want to see, does this work? Does this not work? And when you enable developers to have that kind of power and have that ability to go really fast, there's all kinds of like things that we can enable for the business that help cost savings, better security, all kinds of stuff far beyond just simple, “Hey, here's more features. Here's more features.”

Emily: How easy do you think it is for operations teams to sort of shift to think, like, “Our job is to make things as easy as possible for developers?”

Dave: I don't think it's that hard, actually, mostly because if we're starting to look at things from a DevOps mindset, we're understanding that the whole goal is to optimize the entire system; it's not to optimize a single point in the system. And I always advocate that operations teams report up through the same reporting structure as the engineering teams do. The worst thing you can do is silo it off so all the operations teams report to the COO, and all the engineering teams report to the CTO. Like, that's awful because what you want to do is you want to align the outcome so that everybody's working towards the same goal, and now we can start to partner up together in order to be able to achieve those goals. And so, one of my favorite examples of this from, like, enabling the developers to go fast, and doing that in partnership with operations was, I worked at a company, and we had a storage system, and we were storing all this stuff in a database, and we were paying a lot of money to store all this data for our customers.

And that's what the customers were paying us for, was to store their data. And so the developers had this idea that they wanted to try this other way of storing the data. And so, we worked with them—the operations teams work with them, “How do you want to do this? What kinds of things do you need? What's going to work best? Is this going to work best? Is that going to work best?” And we had a lot of collaboration, and, “Here's where we're going to launch these new things, and we're going to try them out. And this is how we're going to try them out.”

And it wasn't a process that happened overnight: from beginning to end of this project, it probably took, I don't know, a year and a half or something like that, of iterating, and trying, and testing, and making sure it's safe, and all these other stuff. But in the end, we wound up shutting off the old database system and talking to the engineers about what that meant for the business. They said a conservative estimate would be that we saved the company 75 percent on storage costs. That's the conservative estimate. I mean, that's insane, right? 75 percent for their biggest cost. That was the biggest cost of the company, and we knocked it down by 75 percent, at minimum.

And so this idea of enabling this cloud approach of going quickly, and taking advantage of all these resources, and moving fast without impediments, that can have some major impact. And it's not operations teams doing that; it's not development teams doing that; it's operations and development teams doing that together in partnership to achieve those pretty awesome business outcomes.

Emily: In that particular case, who had the initial push? Who had this initial idea that let's figure out a better way to approach storage?

Dave: Well, I mean, we got a challenge from the business. The business said, look at our costs. Look at what we're doing. Are there ways that we can improve this so that we can improve our profitability? And so it was a challenge.

And I think the best thing about that is, it wasn't the business telling us how to do it; it wasn't people saying, here's what you should do. The business is saying, “If this is a problem, how do you solve it?” And then they, kind of, got out of our way and said, “Let the engineers do their engineering.” And I think that was kind of fantastic because the results were exactly what they wanted.

But the business is going to look at the problems from a business perspective, and I think it's important that as engineers, we look at the problems from a business perspective as well. We're not showing up for work to have fun and play with computers. We're showing up at work to achieve an objective. That's why we get paid. If you want to hobby around with your computers, you can hobby around at home, but we're getting paid at work to achieve the goals of the business. And so, that was the way that they were looking at the problem, and that's the way that we wound up looking at the problem. Which is the correct way?

Emily: Do you have any other notable examples that come to mind?

Dave: Yeah, I mean, this idea of cloud and being able to go quickly, we had this one problem with that—actually, with that same database engine, which is hilarious, before we wound up replacing it, where we were upgrading the software from one version to another, and we're making a pretty big jump. And so, we spun up the new version of the software; we loaded the data on, and we started seeing how their performance was. And the performance was terrible. I mean, not just, we would have trouble with it; it was unusable. There was no way we could run the business with that level of performance.

And we're like, “What happened? [laughs]. What did we do here?” And so, we went and looked in GitHub at the differences between the old version of the software, and the new version of software. And there was, like, 5000 commits that had happened between the old version and the new version. And so all we had to do was find out which of those 5000 commits was the problem. [laughs]. Which, that's a daunting task or whatever.

But the operations team got down to it, and we built a bunch of tooling, and we started changing some things and making improvements so that we were able to spin up clusters of this software and run a full test suite to determine whether this problem still existed. And that was something we could do in 20 minutes. And so then we started doing what's called git bisecting, but we started searching in a certain kind of pattern, which I won't get into, for which of these—where was the problem? So, we would look, say in the middle, and then if the problem wasn't there in the middle, then we would look between the middle and the right. If it was there, then we would look between the middle and the left. And we kept doing this bisect, and within two days, we had found the exact commit that had caused the problem. And it was them subtracting, like, a nano from a milli, or something like that.

But going back and talking to the CTO afterwards, I said, “You know, if we hadn't built these tools, and we hadn't had this ability to really iterate super quickly in the Cloud, what would you have done?” And he was like, “I have no idea.” He's like, “Maybe we would have spent a couple of days and then given up, maybe we would have just gone in a completely different direction.” But that ability to be able to work so effectively with these cloud tools, and so easily with these cloud tools, enabled us to do something that the business just would just not have had the opportunity to take advantage of at all. And so that was a major win for being able to have operations teams that think about these problems in a completely different way.

Emily: It sounds like, in this particular company, the engineering teams and the business leaders were fairly well-aligned, and able to communicate pretty well about what the end goals are. How common do you find that is with your clients?

Dave: I don't know. I think it's pretty variable. It depends on the organization. I think that is one of the things that I emphasize when I'm working with my clients is how important that alignment is. I sort of talked about it a little bit earlier, when I said you shouldn't have one group reporting to the CTO and another group reporting to the COO.

But also, it's really important for leadership to be communicating this stuff in the proper way. One of the things I loved most about my experience working at Salesforce was, Marc Benioff was the CEO and he would publish what they call V2MOMs, which is like—oh boy, vision, values, metrics, obstacles, and measures, or something. I don't remember what the last thing was. But he would publish his V2MOM, which was basically his objectives for the next time period, whether it was quarterly, or yearly, I don't really remember. But then what would happen was the people that worked for him would look at his V2MOM, and they would write theirs about what they wanted to get accomplished, but showing how what they were doing was in support of what he wanted. And then the people below them would do the same thing. And the people below them would do the same thing. And what you were able to create was this incredible amount of alignment at a 16,000 person company, which is crazy, up and down the ladder so that everybody understands what they're doing, and how it fits into the larger picture, and what they're doing in support of the goals of the business, and the objectives of the business, and that goes all the way down to the most junior engineer. And I think having that kind of alignment is, I mean, it's obviously incredibly powerful. I mean, Salesforce is a rocket ship and has been for a long, long time. And Google does this for their OKRs, and that there was that thing that was popularized by Intel as well; there's a whole bunch of these things. But that alignment is phenomenal if you want to have a lasting, high performing organization.

Emily: When you see companies that don't have that alignment, or even just, it seems like the engineering team maybe doesn't entirely understand where the business is going, or even the business doesn't understand what the engineering team is doing, what happens and where is the communication going wrong?

Dave: I mean, you see the frustration. You see the fracturing. You see the silos. You see a lot of finger-pointing. I've definitely worked with some clients where the ops team hates the dev team; the dev team hates the ops team. I remember the dev team saying, “Ops doesn't actually want to do any work. They just want to invent stuff for themselves to work on. And that's how they want to spend their day.” And the ops team is saying, “The developers don't even understand anything about what we're doing, and they just want to go o—” you know, there's all these crazy, awful made up stories.

And if you've ever read the book Crucial Conversations—they also have a course, or whatever—one of the things they talk about is you need to establish mutual purpose in order to have a difficult conversation. And I think that's really important for what we're talking about in the business: we need to establish mutual purpose, just we talked about with DevOps, there's only one goal. And the other thing that they say in that class, or in that book, is that when we are going to have a conversation with somebody that we are not getting along with, we invent a story that explains why they behave the way that they do, and every time we see something that validates that story, then it's even more evidence that that story is actually correct. The problem is, is it's a story. Like it's not real. It may seem real to us; it may feel real to us. But it's a story. It's something that we made up. And so, that's the kind of outcomes that you get when you have this fracturing, where you don't have this alignment up and down. You have people telling these stories like, “Operations doesn't want to do any real work. They just want to make stuff up for themselves to work on.” Which is, if you're not in that environment, if you're someone like you and me looking from the outside, that's absurd. But I can certainly see how you can make up a story that gets continually validated by what you see because you're looking for evidence that supports your story. That's part of what makes you think that you're right, is that you're always searching for this evidence. And so obviously, those are not going to be high performing organizations. That's why it's so important to get that kind of alignment.

Emily: Going back to this idea of sort of moving to cloud-native, what do you think are some of the surprises or misconceptions that come up when teams are moving to more cloud-native approaches?

Dave: I feel like my clients generally are not terribly surprised. I think by the time that someone's reaching out to me, they are feeling a lot of pain, and they know that things have to change, and they are looking for what are the ways that things have to change? I don't ever have to go into a client and convince them that they need to do it better. The clients that are coming to me, recognizing that they're having a problem, and so it's really just getting them to stop focusing on what we call an SRE toil, which is popularized by Google, which is—I don't remember the exact definition, but it's basically manual work that's devoid of enduring value, that's repetitive, it's automatable, it's just repeat, repeat, repeat, repeat: we're not making improvements anymore. And so once we start to have this Kaizen mindset, this idea of continual improvement at all times, instead of just trying to keep the business running, that starts to enable all this kinds of stuff.

And that's why we talked about building things in sort of a cloud-native manner. We're talking about that ability to go fast. We're talking about that ability to enable things, and part of that is this idea of continual improvement; this idea of always making things better. A lot of this comes out of Agile as well. I always talk to people about their sprint retrospectives, and I say, “It's your opportunity to make your team better. It's your opportunity to make your environment better. It's your opportunity to make your company better.” And I was like, “The worst thing that you could do in Agile is if in January of last year, and in January of this year, your team is just as good as it was.” That's terrible. Your team needs to be much better than it was. And so enabling developers to go quickly and all that other stuff. And putting all those things in place is a big part of that.

Emily: Anything else that you want to add about this topic that I didn't think to ask?

Dave: I mean, I think embracing these principles is really important. I think that if you look at the companies who are trying to go fast, and don't embrace these principles, these cloud-native ideas, or just even these cloud computing ideas, it basically becomes technical debt that keeps building up, and building up, and building up. And everybody knows accumulating tons of technical debt is not going to help your organization to move faster; it's not going to help you achieve all those great business outcomes that you want to get out of the State of DevOps report. And so I've seen situations where they have not been able to make that transition into this way of looking at the world, and the environment becomes really fragile; it becomes really brittle; it becomes really hard to make changes, and the only way is to make changes is to double down on the technical debt, and accumulate more of it, to the point where eventually they wind up spinning up an entire team whose sole purpose is to try to undo the mess that's been created.

And you don't want that. You don't want to allocate a team to start unpacking your technical debt. You'd rather just not accumulate that technical debt as you're going along. And so I think it's really crucial for businesses that want to be successful in the long term that they start to embrace these ideas early. And obviously, if I'm a startup and I want to embrace a lot of these cloud-native things, that's a lot easier than if I'm a well-established company. I, in my consulting practice, I don't really work with startups because they don't tend to have these problems. They don't tend to accumulate a lot of technical debt because they are founded with this idea of going quickly and being able to empower developers and enable people to go quick.

To your point earlier, the companies that I'm working with are the ones who are making this transition, where they've been running in the data center, or maybe they built an environment in the Cloud, but it's just not operating the way that they expected, and they're paying ridiculous amounts of money [laughs] to run stuff in AWS, where we thought, “Hey, what's going on? This isn't supposed to be this way.” But startups have the ability to do this much easier because they're unencumbered. And then as they grow, and they start to introduce more process because that stuff is inevitable that we're going to need to do that, that's when these things become even more important that we make sure that we're keeping them in mind and we're doubling down on them, and we're not introducing lean waste into the system and stuff like that, that will ultimately catch up with us.

Emily: It's so true. All right, just a couple more questions. What is your favorite engineering tool?

Dave: Ah. I mean, I'm supposed to give some kind of DevOps-y, it's not about the tools answer, but this week I think it was, on Twitter, I saw somebody else put up a SmokePing graph. And most people are not going to have heard of SmokePing. I worked at multiple ISPs in my career already, so the networking stuff is important to me. But wow, I love a SmokePing graph. And it's basically just a bunch of pings that are sent to some target, and then they're graphed when they come back, but instead of saying, “I sent one ping, and I came back with 20 milliseconds,” it sends 20, and then it graphs them all at that time point, so you can actually see density. It's basically before everybody came up with the idea of heat maps, this was one of the original heat map tools, and I still run SmokePing in my house just to see the performance of my home network going out to different parts of the internet, and that's definitely my favorite tool.

Emily: Where can listeners connect with you? Website?

Dave: Yeah, yeah, that's a great question. So, if people are interested in my business, I'm at mangoteque.com, M-A-N-G-O-T-E-Q-U-E. That was a fun name invented by Corey Quinn of The Duckbill Group and I loved it, and so I wound up using it. And they could also find me on LinkedIn obviously, or Twitter at @DaveMangot. M-A-N-G-O-T, and I post a lot on there about things that I've observed. I post a lot on there about DevOps. I post a lot on there about taking a scientific approach to a lot of the things we're doing, not just in terms of the scientific method, but like in terms of cognitive neuroscience, and things like that. And I also write a monthly column for CIO.com.

Emily: Well, Dave, thank you so much for joining me.

Dave: Thank you for having me, Emily, this was really fun.

Announcer: Thank you for listening to The Business of Cloud Native podcast. Keep up with the latest on the podcast at thebusinessofcloudnative.com and subscribe on iTunes, Spotify, Google Podcasts, or wherever fine podcasts are distributed. We'll see you next time.

This has been HumblePod production. Stay humble.

View Details

This conversation covers:

  • How Frame.io was faced with the decision to be cloud native or cloud-enabled — and the business and technical reasons why Frame.io chose to be cloud native.
  • How Abhinav successfully built a world class cloud-native security program from the ground up to protect Frame.io users’ sensitive video content. Abhinav also talks about the special security considerations for truly cloud native applications.
  • Cloud native as a “journey without a destination.” In other words, there is no end point with cloud native transitions, because new technologies are always being developed.
  • Why Abhinav is a firm believer in both ISEs and GitOps, and why he thinks the industry should embrace both of these strategies.
  • The challenge of not only maintaining security in this type of environment, but also communicating security issues to various stakeholders with different priorities. Abinhav also talks about the role that specialists like AWS and machine learning experts can play in furthering security agendas.
  • Common misconceptions about cloud native security.
  • Frame.io’s decision to roll out Kubernetes, and why they are also considering adding chaos engineering to fortify against unexpected issues.
  • Tool and vendor overload, and the importance of trying to find the right tools that fit your infrastructure.

Links:

  • Frame.io: https://frame.io/
  • Connect with Abhinav on LinkedIn: https://www.linkedin.com/in/absri/
  • The Business of Cloud Native: http://thebusinessofcloudnative.com

Transcript

Announcer: Welcome to The Business of Cloud Native podcast where we explore how end users talk and think about the transition to Kubernetes and cloud-native architectures.

Emily: Welcome to The Business of Cloud Native. I'm Emily Omier, your host, and today I am chatting with Abhinav Srivastava. Abhinav, can you go ahead and introduce yourself and tell us about where you work, and what you do.

Abhinav: Thanks for having me, Emily. Hello, everyone. My name is Avinash Srivastava. I'm a VP and the head of information security and infrastructure at Frame.io. At Frame, I am building the security and infrastructure programs from ground up, making sure that we are secured and compliant, and our services are available and reliable. Before joining Frame.io, I spent a number of years in AT&T Research. There I worked on various cloud and security technologies, wrote numerous research papers, and filed patents. And before joining AT&T, I spent five great years in Georgia Tech on a Ph.D. in computer science. My dissertation was on cloud and virtualization security.

Emily: And what do you do? What does an average day look like?

Abhinav: Right. So, just to tell you where I answer the question where I work: so I work at Frame.io, and Frame.io is a cloud-based video review and collaboration startup that allows users to securely upload their video contents to our platform, and then invite teams and clients to collaborate on those uploaded assets. We are essentially building the video cloud, so you can think of us as a GitHub for videos.

What I do when I get to office—apart from getting my morning coffee—as soon as I arrive at my desk, I check my calendar to see how's my day looking; I check my emails and slack messages. We use slack primarily within the company doing for communication. And then I do my daily standup with my teams. We follow a two-week sprint across all departments that I oversee. So, a standup gives me a good picture on the current priorities and any blockers.

Emily: Tell me a little bit about the cloud-native journey at Frame.io? How did the company get started with containers, and what are you using to orchestrate now? How have you moved along in the cloud-native journey?

Abhinav: We are born in the cloud, kind of, company. So, we are hosted in Amazon AWS since day one. So, we are in the cloud from the get-go. And once you are in the cloud, it is hard not to use tools and technologies that are offered, because our goal has always been to build secure, reliable, and available infrastructure. So, we were very, very mindful from the get-go that while we are in the cloud, we can choose to be cloud-native or just cloud-enabled. Means use tools, just virtual machines, or heavyweight virtual machines, and not to use container and just host our entire workload within that.

But we chose to be cloud-native because, again, they wanted to boot up or spin up new containers very fast. As a platform we, as I mentioned, we allow users to upload videos, and once the videos are uploaded, we have to transcode those videos to generate different low-resolution videos. And that use case fits with the lightweight container model. So, from the get-go, we started using containerized microservices; orchestration layer; From AWS, their auto-scaling; automation infrastructure as a code; monitoring. so all those things were, kind of, no brainer for us to use because given our use case and given the way we wanted to be a very fast uploader and transcoder for all of our customers.

Emily: This actually leads me to another question: have you guys seen a lot of scaling recently as a result of stay-at-home orders and work from home?

Abhinav: Right. So, we are seeing a lot more people moving towards remote collaboration tools who are actually working in the production house since they have to work from home now. So, they are now moving to these kind of tools such as Frame.io. And we do see a lot more customers joining our platform because of that. From the traffic perspective, we did not see much increase in the web traffic or load our infrastructure, because we have always set up the auto-scaling and our infrastructure can always meet these peak demands. So, we didn't see any adverse effect on our infrastructure from these remote situations.

Emily: What were some of the other advantages? Like you were talking about that you had the choice to be either cloud-enabled or truly cloud-native? What were the biggest, you know—and I'm interested, obviously in business rationale to the extent you can talk about it—for being truly cloud-native?

Abhinav: So, from business perspective, again, a goal was to [basic] secure available and reliable production infrastructure to offer Frame.io services. But cloud-native actually helped us to faster time to market because our developers are just focusing on the business logic, deploying code. They were not worried about the infrastructure aspect, which is good. Then we’re rolling out bug fixes very quickly through CI/CD platform, so that, again, we offer the better [good] services to our customer.

Cloud-native helped us to meet our SLA and uptime so that our customer can access their content whenever they would like to. It also helped us securing our infrastructure and services, and our cost also went down because we were scaling up and down based on the peak demand, and we don't have to provide dedicated resources, so that's good there. And it also allowed us to faster onboard developers to our platform because we are using a lot of open source technologies, and so the developers can learn quickly—there are a lot more resources out there for them to learn. And it also helped us avoid vendor lock-in. We are relying on more and more open-source projects, CNCF [unintelligible] projects, so that has helped us. And more importantly, it is helping us stay competitive because in this industry—in this time—we would like to be available, we would like to be secure. So, for our customers to stay doing their job that they used to do in an office setting or in a non-remote setting, and we can continue providing help that they need.

Emily: How has this changed the security story?

Abhinav: So, obviously, security story is same what we have before because, I mean, we allow people to have upload their media content to our platform. So, that's very sensitive content. So, we always wanted to make sure that they stay secure. And for that, we have built a world-class security program from ground up, with emphasis on product security, cloud security, security data science, and also compliance and privacy program. So, we are doing what we used to do: making sure that content is still secure, our infrastructure follows the AWS security best practices, we can identify vulnerability within our application and fix it. So, again, as I said, that it hasn't changed much from security perspective, as far as Frame.io’s daily operations are concerned.

Emily: How does having a truly cloud-native application, how is that different from a security perspective from something that isn't cloud-native?

Abhinav: So, security is very important whether you are cloud-enabled or cloud-native. So, security is very important for all the services. Being in the world of microservices and in the container, actually, it helped us to model the application behavior. For example, if you have one very big monolithic application, it does so many things, so it's really hard for you to know to find out what's the normal execution pattern. And when this application is going to—if it attacked, how it's going to behave, how is abnormal execution look like? But in the microservices world, since each application, each microservices is getting one job. So, you can create a good model of behavior of that container.

Or even if you are monitoring their runtime behavior, you know that what kind of processes are going to be invoked from that container? What kind of network connections are going to be made? What are the files are going to be accessed by the services within the host, or within S3, or other resources? So, you know their interaction pattern—execution pattern, and that, you can qualify, both in terms of your security rules that you want to create on the infrastructure for those services, or you can create a better anomaly detection or machine learning models for those behavior. And we did both in our infrastructure to keep them secure.

Emily: And how do conversations about security go when you talk with different stakeholders. I'm curious to know if there's any sort of miscommunications, or things that are lost in translation when you're talking about security with, say, the development team; with the business stakeholders; with platform engineers. What are some of the things—anything that gets lost in translation?

Abhinav: So, there are two parts of this question. In general, having a discussion around cloud-native services and the security of cloud-native services. Because there are various ways you can deploy a service in the cloud, you can have a service deployed in the cloud just by running a bunch of VMs, or you can deploy it using cloud-native architecture where you have doing all those things. But the cloud-native architecture requires you to think of all the stages of the services. For example, how will SLAs, SLOs, SLIs look like for this service? Or, how do you monitor the service when it execute? How will you protect these services when you deploy them? What kind of resources are going to be accessed by this service? How will create their identity and management rules there? How would you deploy it and how would you create network rules for that so that you can do it in a principle of least privileged fashion, you can execute these services?

So, you need to do proper planning that how would a new service going to interact with other services in the infrastructure. And these non-functional requirements are, many times, described poorly or not written at all because as a developer, you would like to create service and deploy service, and so that customer can use it. And these are the things behind the scenes we have to think about it. And we, as a team are working very actively to bridge this knowledge and semantic gap so that these things don't get lost in the translation when you're thinking about the service.

Emily: What about when you talk to say, business stakeholders? Is there anything that gets lost in the translation?

Abhinav: So, I mean, in the business sense, we always have to keep the discussion at a very high level. That, what's a use of service? Or, where we should deploy? Who are going to be the users? So, at that time, we don't want to talk about those underlying infrastructure-related issues because at the business level, we would like to know that how the service is going to function, and mostly functional requirements. But at the low level, we would like to think about that when we are about design these services, what are the things we have to worry about in order for that service to deploy securely and reliably?

Emily: How important is security to Frame.io? Not every company thinks the same about security, I should say.

Abhinav: And that's a great question. I think for us, security is very important. I know every company says that, but I think we truly mean that. So, we are close to 150 employees, but I was hired around when I was a [00:12:31 unintelligible] employee as a head of security. So, that shows that we care about security. And I have been building security from ground up. We got our SOC 2 Type II compliance when we was around 70 employees. And there are companies out there who are doing SOC 2, and they are thousand employees. So, we are GDPR compliant; we are working towards our CCPA compliance, and we are TPN compliant as well. TPN stand for Trusted Partner Network, which is the [same world] media, and entertainment companies, and industry users. And we were the first few companies who got that certification, also. So, we care about security very much because we allow users to upload their contents in our cloud and we make sure that those contents remain secure.

Emily: And so, is there any tension that you feel between talking about security or making things as secure as possible, and either business stakeholders or other parts of the IT team?

Abhinav: So, there is definitely attention. [laughs]. If I say no, then I would be lying because our goal—engineers or developers or service creators, they want to deploy the service. They will get satisfaction once the customer start using those services. And our job is to make sure to—we put some guardrails in place—or barriers in place so that we can vet the application, we can vet the service, we can do the proper testing, we can make sure that by deploying the service, we don't increase our exploitable surface.

So, that kind of tension will always be there because, by nature, security's job is to make sure that whatever is deployed is secure. Our infrastructure is secure and the service owner’s job is to deploy the service. But I think what we are trying to do in the organization, we are trying to take a risk-based approach because security is just another business function. The way sales is important, the way engineering is important, the same way that security is important. And there's a risk in this environment of not meeting sales targets, same way there's a risk of getting breached.

So, how do we provide a risk-based methodology so that when we talk about security, we talk in terms of risk; we talk in terms of probabilities versus possibilities? Because there is always possibility of something going wrong, but what's the probability of something happening? And that basically gives us some way of talking to other business-holders saying that, “Okay, if you deploy the service the risk is high. But the risk is high because the likelihood of getting breached is high, but impact would be very low. So, since risk is the product of impact and likelihood, overall the risk is low.” But sometimes the risk is that chance of getting attacked is very low, but the impact could be very high. Again, you will have risk low because probability of actually happening that event is low.

So, that basically gives us some common language we can use to talk to other business-holders because risk is being used as a language across other departments. We try to use the same language to convey cybersecurity risk as well.

Emily: Since starting with Frame.io and building this security program from the ground up, what surprises have you encountered?

Abhinav: I would say there were many surprises. First of all, I had those surprises because I come from a background from research and development. There, goal was to develop services, goal was to think about new security product, and goal was to think of attack and coming up with defenses for them. Having the responsibility of building the security program from ground up, or having to adjust this risk-based mentality was a big surprise because it's not that just because there is a bug, engineering is going to fix it. You have to show the impact of that bug. You should have a proper [unintelligible] associated with that. You have to show that what are the ways that bug can be launched. So, it means, just because you care about security, doesn't mean that everybody else cares about security. So, you have to keep the communication on. You have to always talking, you have to always adjusting, and you have to use the right language to the right person that you are talking to.

Emily: What tips do you have about adjusting your language for different audiences and getting them to understand what you're talking about?

Abhinav: So, one thing is to use risk-based methodology. That is saying that, “Oh, we have a bug, or we have a high priority bug.” I think saying that, “What is the impact of that bug? How would that bug be exploited in a real setting?” I think those things are important because people care about security, but then they have hundred other things to do, as well. So, how do you talk to their language?

And also building the right team, as well. So, if you want to target product security, you have to have a product security specialist, who can understand these nuances; who can understand what are the different attacks. Some companies build a security team with many generalists. I took an approach where I'm building team with the specialists.

So, for product security, I have two core product security engineers who have done this thing many times before. For cloud security, I have a specialist who knows about AWS Cloud and everything. For security data science, I have a machine learning expert. So, for each of those roles that you have mined, you try to fill the position with the right set of people. And coming back to this cloud-native security.

I think one thing is very important in the cloud-native world, as I have realized lately, that infrastructure as a goal is very important piece for securing your cloud. It's not that I or the team don’t know about it, but the temptation to do things quickly sometimes resorting to manual work instead of writing your Terraform or CloudFormation. So, you can do things quickly, but then the chances of you making error are also high. Because if you go to Terraform, you can follow the regular CI/CD process, you can have your pull request approved by somebody, and chances of finding a error quickly is high.

And for security purposes, infrastructure code is a blessing. Because you can put proper guard rails in place to make sure that nobody does manual operation in the infrastructure, and everything goes through proper approval process, and that will—as a head of security if you know that if somebody wants to do anything or open any port in the infrastructure, two people are going to look at it and then they're going to have a dialogue with each other, and they’re going to find out the real need for opening that port. Your life will be a lot simpler.

Emily: What do you think are some misconceptions about cloud-native security, both inside the engineering department—so developers, for example—and then outside in the rest of the company?

Abhinav: I think misconception that I view—and it's my opinion—is that the only thing that is important is deploying fast, or moving to production very fast. I think there are so many things has to be done behind the scene in order for you to move fast. And if you don't do those things, then it means that either you're going to break your application, or you're going to make your infrastructure insecure. So, for example, if you have a CI/CD set up and you want to deploy a business logic, and you think that, “Oh, I can code that thing in AWS Lambda functions.” AWS Lambda function is completely managed service. You went ahead and coded in Python, and your service is up and running. But now in doing so, what you did quickly that you forgot to follow the best practices that Lambda function has to be within the VPC; you need to generate an IAM role that has restricted permission; you have to make sure that proper security groups has to be attached to Lambda functions so that it is not open to www. And those things are part of misconception that, “Oh, if I have to do something, AWS allows that we can do it quickly.” That's what we are trying to do. We are trying to come up with a set of best practices for each of those resources as a team, writing documents, sharing with engineering that, “Okay, you want to do it? Sure, go ahead, do it, but just follow these best practices.” So, that even if you SAM or Terraform, whatever you want to use to deploy your application, make sure that best practices are always followed.

Emily: Can you think of any misconceptions about cloud-native security that, say, somebody might have if they're coming from a legacy environment: managing security but in a very different type of environment.

Abhinav: So, I mean, cloud-native security is all about making sure that your microservices are secure, the kind of access pattern they have, kind of network pattern they have. So, I think one misconception is that—you can think of misconception is, if you are coming from a monolithic world, where you have logged on your services, but just by assuming that you have a parameter between outside world and inside world, so your firewall rules are just like that between in and out. But that parameter is blurred now. There is no such thing as a “them versus us.” It's all blurred now.

So, in the microservices world, instead of North/South traffic going up and down. You have to think about East/West traffic as well. So, making sure that your service communication are secure as well: you make sure you use proper cryptography, make sure your endpoints are authenticated so that your services are not compromised. Because if one service compromised, if you don't use proper control among those services, then your other services can be compromised very quickly. And that's the problem when we go from monolithic application to microservices.

Emily: Do you think that people outside of the security team understand that distinction?

Abhinav: I would say they do, to the extent that they know about it, but then when we have to actually implement it, there are always some concerns that it is going to slow down our application, it is going to introduce latency in the application. So, people do understand that okay parameter is going away, but to the extent that they know about it, but when you—again, when we start implementing it, there is always concern that how it's going to play out.

Emily: Do you think Frame.io is fully cloud-native? Do you think there's anything that you could do to be more quote-unquote, “cloud native.”

Abhinav: So, in my opinion, it is a journey without any destination. Just like security, you can never say, “I’m secure.” You will have to adjust your control based on the threats or attacks going on. In the same way, there is no end to transition to cloud-native because new technologies are coming, and we will have to evaluate new tools that can help us realize our business goals effectively. So, we are cloud-native, but still, we can do a lot more things, given time and resources.

So, in some concrete world that we are doing right now, that we are creating more tools for developers to perform tasks themselves. So, creating more self-serve culture. As I said that moving towards more [IFC] model, and so on. And for that, we are setting up guardrails so that they can perform those operations within those boundaries without impacting security and reliability. We are also looking into ways to extend Kubernetes. Because Kubernetes is in itself a full cloud platform with a lot of possibilities. So, we are interested in making it more programmable for our environment. But these are ongoing things that we'll have to continue doing it.

Emily: Do you have any other next steps that you could share? What's next in your journey?

Abhinav: So, we rolled out Kubernetes in our infrastructure last December, and that move paid us off. So, we are building more tools on Kubernetes. As I said, that we are going towards more self-service style of architecture where developers can do a lot more things within those guardrails and we are also looking into ways to introduce chaos engineering in our environment because we do things fast, but we break things fast as well. [laughs]. So, one small configuration error can create severity zero alert. So, what we need is a good chaos engineering practices to simulate these areas, so that everybody can train on these events and know how to prevent and respond to such problems. That will reduce our incident resolution time as well.

Emily: When—sort of last question: anything else that you would like to add?

Abhinav: Two things, I think. One thing is we all should be going towards IFC and GitOps; infrastructure code and GitOps. If this is the one takeaway from this podcast, is that that's the way to go. I know manually doing work is tempting, but that creates problem down the road. So, life will be a lot simpler if we go with the IFC and GitOps.

Second thing is that I feel this pain, and many other people are facing the same way, that there are too many tools and vendors out there. So, it's really hard to choose from what is going to work in your environment. CNCF is helping us by highlighting some of these projects by assigning proper maturity levels, like sandbox incubation, and graduated project, so on, but it still is very challenging to find the right tooling that fits your infrastructure. So, always make sure that when you choose a new technology, see how it's going to be working with your existing technologies because it's not that easy to throw away an existing thing because all these things that the tool that you try, it also complicates your security as well because you just do not know how it's going to play out when you deploy this new technology in your environment where the other tools and services are running. So, I think we have to evaluate all tools carefully to make sure that we understand its a security and reliability impact on our existing infrastructure.

Emily: What is your can't live without engineering tool or security tool?

Abhinav: Huh, that's a good question. Right now, one tool that I cannot live without is Falco. That is a runtime container monitoring solution. We invested a lot on it, and it is paying off in terms of the kind of alert it is generating, kind of visibility it is providing in our infrastructure. And one tool I can't leave off from both from security infrastructure perspective is Slack because we have done a lot of automation to bring all these alerts through Slack. So, all of our ops happen via Slack. So, I think these are the two tools I’m relying a lot in terms of visibility and in terms of response.

Emily: Well, thank you so much for joining me.

Announcer: Thank you for listening to The Business of Cloud Native podcast. Keep up with the latest on the podcast at thebusinessofcloudnative.com and subscribe on iTunes, Spotify, Google Podcasts, or wherever fine podcasts are distributed. We'll see you next time.

This has been HumblePod production. Stay humble.

View Details

In this episode of the Business Cloud Native, host Emily Omier talks with Jon Tirsen, who is engineering lead for storage at Cash App. This conversation focuses on Cash App’s cloud native journey, and how they are working to build an application that is more scalable, flexible, and easier to manage.

The conversation covers:

  • How the need for hybrid cloud services and uniform program models led Cash App to Kubernetes.
  • Some of the major scaling issues that Cash App was facing. For example, the company needed to increase user capacity, and add new product lines.
  • The process of trying to scale Cash App’s MySQL database, and the decision to split up their dataset into smaller parts that could run on different databases.
  • Cash App’s monolithic application, which contains hundreds of thousands of lines of code — and why it’s becoming increasingly difficult to manage and grow.
  • How Jon’s team is trying to balance product/ business and technical needs, and deliver value while rearchitecting their system to scale their operations.
  • Why Cash App is working to build small, product-oriented teams, and a system where products can be executed and deployed at their own pace through the cloud. Jon also discusses some of the challenges that are preventing this from happening.
  • How Cash App was able to help during the pandemic, by facilitating easy stimulus transfers through their service — and why it wouldn’t have been possible without a cloud native architecture.

Links:

  • Cash App: https://cash.app/
  • Square: https://squareup.com/us/en
  • Jon on Twitter: https://twitter.com/tirsen?lang=en
  • Connect with Jon on LinkedIn: https://www.linkedin.com/in/tirsen/?originalSubdomain=au
  • The Business of Cloud Native: http://thebusinessofcloudnative.com

Transcript

Announcer: Welcome to The Business of Cloud Native podcast where we explore how end users talk and think about the transition to Kubernetes and cloud-native architectures.

Emily: Welcome to The Business of Cloud Native. My name is Emily Omier, I'm here chatting with Jon Tirsen.

Jon: Happy to be here. My name is, as you said, Jon Tirsen, and I work as the engineering lead of storage here at Cash App. I've been at Cash for maybe four or five years now. So, I've been with it from the very early days. And before Cash, I was doing a startup, that failed, for five years. So, it's a travel guide in the mobile phone startup. And before that, I was at Google working on another failed product called the Google Wave, which you might remember, and before that, it was a company called ThoughtWorks, which some of you probably know about as well.

Emily: And in case people don't know, the Cash App is part of Square, right?

Jon: Yes. Cash App is where we're separating all the different products quite a lot these days. So, it used to be called just Square Cash, but now it has its own branding and its own identity, and its own leadership, and everything. So, we're trying to call it an ecosystem of startups. So, each product line can run its business the way it wants to, to a large degree.

Emily: And so, what do you actually spend your day doing?

Jon: Most of my days, I'm still code, and doing various operational tasks, and setting up systems, and testing, and that sort of thing. I also, maybe about half my day, I spend on more management tasks, which is reviewing documents, writing documents, and talking to people trying to figure out our strategy and so on. So, maybe about half my time, I do real technical things, and then the other half I do more management stuff.

Emily: Where would you say the cloud-native journey started for you?

Jon: Well, so a lot of Square used to run on-premises. So, we had our own data centers and things. But especially for Cash App, since we've grown so quickly, it started getting slightly out of control. We were basically outgrowing—we could not physically put more machines into our data centers. So, we've started moving a lot of our services over to Amazon in this case, and we want to have a shared way of building services that would work both in the Cloud and also in our data centers.

So, something like Kubernetes and all the tools around that would give us a more uniform programming model that we could use to deploy apps in both of these environments. We started that, two, three years ago. We started looking at moving our workload out of our data centers.

Emily: What were the issues that you were encountering? Give me a little bit more details about the scaling issues that we were talking about.

Jon: There two dimensions that we needed to scale out the Cash App, sort of, system slash [unintelligible] architecture. So, one thing was that we just grew so quickly that we needed to be able to increase capacity. So, that was across the board. So, from databases to application servers, and bandwidth, everywhere. We need to just be able to increase our capacity of handling more users, but also we were trying to grow our product as well. So, at the same time, we also want to build and be able to add new features at an increased pace. So, we want to be able to add new product lines in the Cash App.

So, for example, we built the Cash Card, which is a way you can keep your money in the Cash App bank accounts, and then you can spend that money using a separate card, and then we add a new functionality around that card, and so on. So, we also needed to be able to scale out the team to be able to have more people working on the team to build new products for our users, for our customers. Those are the two dimensions: we needed to scale out the system, but we also needed to have more people be able to work productively. So, that's why we started trying to chop up—we have this big monolith as most companies probably do, which that's I don't know how many hundreds of thousands of lines of code in there. But we also wanted to move things out of that, to be able to have more people contribute productively.

Emily: And where are you in that process?

Jon: Well, [laughs], we're probably adding still adding code at an exponential rate to the monolith. We're also adding code at an exponential rate outside of the monolith, but it just feels so much easier to just build some code in the monolith than it is outside of it, unfortunately, which something we're trying to fix, but it's very hard. And it is getting a little bit out of hand, this monolith now. So, we have, sort of, a moratorium on adding new code to the monolith now, and I'm not sure how much of an effect that has made. But the monolith is still growing, as well as our non-monolith services as well, of course.

Emily: When you were faced with this scaling issue, what were the conversations happening between the technical side and the business owners? And how is this decision made about the best way to solve this problem is x, is the Cloud, is cloud-native architecture?

Jon: I think the business side—the product owners, product managers—they trust us to make the right decision. So, it was largely a decision made on the technical side. They do still want us to build functionality, and to add new features, and fix bugs, and so on. So, they want us to do that, but they don't really have strong influence on the technical choices we've made. I think that's something we have to balance out.

So, how can we keep on giving the product side and the business side what they need? So, to keep on delivering value to them while we try to rearchitect our system so that we can scale out our operations on our side. So, it's a very tricky balance to find there. And I think so far, maybe we've erred on the side of keep on delivering functionality, and maybe we need to do more on the rearchitecting things. But yeah, that's always a constant rebalancing act we're always dealing with.

Emily: Do you think that you have gotten the increased scalability? How far along are you on reaching the goals that you originally had?

Jon: I think we have a pretty scalable system now, in terms of the amount of customers we can service. So, we can add capacity. If we can keep on adding hardware to it, we can grow very far. We've actually noticed that the last few weeks, we've had an almost unprecedented growth, especially with the Coronavirus crisis. Every single day, it's almost a record.

I mean, there's still issues, of course, and we're constantly trying to stay on top of that growth, but we have a reasonably good architecture there. What I think is probably our larger problem is the other side, so the human side. As I said, we are still adding code to this monolith, which is getting completely out of hand to work with. And we're not growing our smaller services fast enough. It's probably time to spend more effort on rearchitecting that side of things as well.

Emily: What are some of the organizational, or people challenges that you've run into?

Jon: Yeah. So, we want to build smaller teams oriented around products. We see ourselves more of a platform on products these days: we’re not just a single product. And we want to build smaller teams. That is, maybe we have one team that is around our card, and one team around our [unintelligible] trading and so on. And we want to have the smaller teams, and we want them to be able to execute independently.

So, we want to be able to put together a cross-functional team of some engineers, and some UX people, and some product people, and some business people, and then they should be able to execute independently and have their own services running in our cloud infrastructure, and not have to coordinate too much with all of the other teams that are also trying to execute independently. So, each product can do its own thing, and own their own services, and deploy at their own pace, and so on. That's what we're trying to achieve, but as long as they still have to do a lot of work inside of our big monolith, then they can't really execute independently. So, one team might build something that actually causes issues with another team’s products, and so on, and that becomes very complicated to deal with. So, we tried to move away from that, and move towards a model where a team has a couple of services that they own, and they can do most of their work inside of those services.

Emily: What do you think is preventing you from being farther along than you are? Farther along towards this idea of teams being totally self-sufficient?

Jon: Yeah, I think it's the million-dollar question, really. Why are we still seeing exponential growth in code size in our monolith, and not in our services? And I think it's a combination of many, many things. One thing I think, we don't have all of the infrastructure available to us in our cloud, in our smaller services. So, say you want to build a little feature, you want to add a little button that does something, and if you want to do that inside our monolith, that might take you two, three days. Whereas if you want to pull up a completely new service—I think we've solved it at an infrastructural layer, it's very quick and easy to just pull up a new service, and have it run, and be able to take traffic, and so on—but it's more of the domain-specific infrastructures of being able to access all the different data sets that you need to be able to access, and be able to shift information back to the mobile device.

And all these things, it's very easy to do inside a monolith, but it's much harder to do outside of the monolith. So, we have to replicate a big set of what we call product platforms. So, instead of infrastructural platform is more product specific platform features like customer information, and be able to send information back to the client, and so on. And all those things have to be rebuilt for cloud services. We haven't really gotten all the way there yet.

Emily: If I understood correctly from the case study with the CNCF, you sort of started the cloud-native journey with your databases.

Jon: Yes, that was the thing that was on fire. Cash App was initially built as a hack week project, and it was never really designed to scale. So, it was just running on a single MySQL database for a really long time. And we actually literally put a piece of hardware on fire with that database. We managed to roll it, roll it off, of course, didn't take down our service, but it was actually smoking in our [laughs] data centers. It melted the service around it in its chassis. So, that was a big problem, and we needed to solve that very quickly. So, that's where we started.

Emily: Could you actually go into that just a little bit more? I read the case study, but probably most listeners haven't. Why was the database such a big problem? And how did you solve it?

Jon: Yeah, as I said, so we only had a single MySQL database. And as most people know, it's very hard to keep on scaling that, so we bought more and more expensive hardware. And since we were a mobile app, we don't get all the benefits from caching and replica reads, so most of the time, the user is actually accessing data that is already on the device, so they don't actually make any calls out to our back end to read the data. Usually, you scale out a database by adding replicas, and caching, and that sort of stuff, but that wasn't our bottleneck. Our bottleneck was that we simply could not write to the database, we couldn’t update the database fast enough, or with enough capacity.

So, we needed to shard it, and split up the data set into smaller parts that we could run on separate databases. And we used the thing called Vitess for that, which is a Cloud Native Foundation member, a product and [unintelligible] CNCF. And with Vitess, we were able to split up the database into smaller parts. It was quite a large project, and especially back then, Vitess was—it was quite early days. So, the Vitess was used to scale out YouTube and then it was open-sourced. And then, we started using it. I think, not long after that, it was also used by Slack.

So now, currently Slack uses it for most of its data. And we started using it very early, so it was still kind of early days, and we had to build a lot of new functionality in there, and we had to port [00:15:20 unintelligible] make sure all of our queries worked with the Vitess. But then we were able to do shard splitting. So, without having to restart or have downtime in our app, we could split up the database into smaller parts, and then the Vitess would handle the routing of queries, and so on.

Emily: If at all, how did that serve as the gateway to then starting to think about changing more of the application, or moving more into services as opposed to a monolith?

Jon: Yeah, I think that was kind of orthogonal in some ways. So, while we scaled out the database layer, we also realized that we needed to scale out the human side of it. So, we have multiple teams being able to work independently. And that is something we haven't I think we haven't really gotten to completely, yet. So, while we've scaled out the database layer, we're not quite there from the human side of things.

Emily: Why is it important to scale any of this out? I understand the database, but why is it important to get the scaling for the teams?

Jon: Yeah, I mean, it's a very competitive space, what we're trying to do. We have a very formidable competitors, both from other apps and also from the big banks, and for us to be able to keep on delivering new features for our customers at a high pace, and be able to change those features to react to changing customer demands or, like during this crisis we are in now, and being able to respond to what our competitors are doing. I mean, that just makes us a more effective business. And we don't always know when we start a new product line where it's exactly going to lead us, we sort of look at what our customers are using it and where that takes us, and being able to respond to that quickly, that's something that is very hard if you have a big monolith that has a million lines of code and takes you several hours to compile, then it’s going to be very hard for you to deliver functionality and make changes to functionality in a good time.

Emily: Can you think of any examples where you're able to respond really quickly to something like this current crisis in a way that wouldn't have been possible with the old models?

Jon: I don't actually know the details here. I live currently in Australia, so I don't know. But the US government is handing out these checks, right? So, you get some kind of a subsidy. And apparently, they were going to mail those out to a lot of people, but we actually stepped up and said, look, you can just Cash App them out to people. So, people sign up for a Cash App account, and then they can receive their subsidies directly into the Cash App accounts, or into their bank accounts via our payment rails. And we were able to execute on that very quickly, and I think we are now an official way to get that subsidy from the US government. So, that's something that we probably wouldn't have been able to do unless we've invested more to be able to respond to that so quickly, within just weeks, I think.

Emily: And as Cash App has moved to increasingly service-oriented architectures and increasingly cloud-native, what has been surprisingly easy?

Jon: Surprisingly easy. I don't think I've been surprised by anything being easy, to my recollection. I think most things have been surprisingly hard. [laughs]. I think we are still somewhat in the early days of this infrastructure, and there are so many issues; there's so many bugs; there's so many unknowns. And when you start digging into things, it just surprises you how hard.

So, I work in the infrastructure team, and we try to provide a curated experience for our product teams, the product engineering teams, so we deal with that pain directly where we have to figure out how all these products work together, and how to build functionality on top of them. I think we deal with that pain for our product engineers. But of course, they are also running into things all the time. So, no, it is surprisingly hard sometimes, but it's all right.

Emily: What do you think has been surprisingly challenging, unexpectedly challenging?

Jon: Maybe I shouldn't be, but I am somewhat surprised how immature things still are. Just as an example, how hard it is, if you run a pod, in a EKS—Amazon Kubernetes cluster, and you just want to authenticate to be able to use other Amazon products like Dynamo, or S3, or something, this is still something that is incredibly hard to do. So, you would think that just having two products from the same vendor inside of the same ecosystem, you would think that that would be a no-brainer: that they would just work together, but no. I think we'll figure it out eventually, but currently, it's still a lot of work to get things to play well together.

Emily: If you had a top-three wish list of things for the community to address, what do you think they would be?

Jon: Yeah, I guess the out-of-the-box experience with all of these tools, so that they just work together really well, without having to manually set up a lot of different things, that'd be nice. I think I also, maybe this all exists, we haven't integrated all these tools, but something that struck me the other day, I was debugging some production issue—it wasn’t a major issue, but it was an issue that had been an ongoing thing for two weeks—and I just wanted to see what change happened those two weeks ago. What was the delta? What made that change happen? And being able to get that information out of Kubernetes and Amazon—and maybe there's some audit logging tools and all this stuff, but it's not entirely clear how to use them, or how to turn them on, and so on. So, that's a really nice, user friendly, and easy to use kind of auditing, and audit trail tools would be really nice.

So, that's one wish, I guess, in general: having a curated experience. So, if you start from scratch, and you want to get all of the best practice tools, and you want to get all the functionality out of a cloud infrastructure, there's still a lot of choices to make, and there's a lot of different tools that you need to set up to make them work together, Prometheus, and Grafana, and Kubernetes, and so on. And having a curated out-of-the-box experience that just makes everything work, and you don't have to think about everything, that would be quite nice. So, Kubernetes operators are great, and these CRDs, this metadata you can store and work with inside of Kubernetes is great, but unfortunately they don't play well with the rest of the cloud infrastructure at Amazon, at AWS.

Amazon was working on this Amazon operator, which you would be able to configure other AWS resources from inside of the Kubernetes cluster. So, you could have a CRD for an S3 bucket, so you wouldn't need a Terraform. So right now, you can have Helm Charts and similar to manage the Kubernetes side of things, but then you also need Terraform stuff to manage the AWS side of things, but just something thing that unifies this, so you can have a single place for all your infrastructural metadata. That would be nice. And Amazon is working on this, and they open-sourced something like an AWS operator, but I think they actually withdrew it and they are doing something closed-source. I don't know where that project is going. But that would be really nice.

Emily: Go back again to this idea of the business of cloud-native. To what extent do you have to talk about this with business stakeholders? What are those conversations look like?

Jon: A Cash App, we usually do not pull in product and business people in these conversations, I think, except when it comes to cost [laughs] and budgeting. But they think more in terms of features and being able to deliver and have teams be able to execute independently, and so on. And our hope is that we can construct an infrastructure that provides these capabilities to our business side. So, it’s almost like a black box. They don't know what's inside. We are responsible for figuring out how to give it to them, but they don't always know exactly what's inside of the box.

Emily: Excellent. The last question is if there's an engineering tool you can't live without?

Jon: I would say all of the JetBrains IDEs for development. I've been using those for maybe 20 years, and they keep on delivering new tools, and I just love them all.

Emily: Well, thank you so much for joining.

Jon: Thanks for inviting me to speak on the podcast.

Announcer: Thank you for listening to The Business of Cloud Native podcast. Keep up with the latest on the podcast at thebusinessofcloudnative.com and subscribe on iTunes, Spotify, Google Podcasts, or wherever fine podcasts are distributed. We'll see you next time.

This has been HumblePod production. Stay humble.

View Details

Emily and Dejan cover the following points:

  • 8x8’s journey to a leading cloud technology provider.
  • Why 8x8 decided to migrate to Kubernetes, a move that gave them the flexibility to run workloads wherever they want.
  • Dejan’s thoughts on the Kubernetes migration, and how it’s helped the company improve its operations. For example, Kubernetes has helped 8x8 migrate away from several legacy systems.
  • The biggest challenges and surprises that the 8x8 team experienced during their migration journey, such as getting engineering teams to embrace a culture built around monitoring, observability, and documentation.
  • How 8x8 has avoided “feature bloat” and maintained a product that performs at a high level, while staying true to the features that are important for its core customer base.
  • The strategy of obtaining buy-in from stakeholders and fellow executives by focusing on business problems, instead of technical issues. This included cost, velocity of innovation, global scale, and so on.
  • How 8x8’s cloud-native architecture has made it faster and easier to scale.

Transcript

Announcer: Welcome to The Business of Cloud Native podcast where we explore how end users talk and think about the transition to Kubernetes and cloud-native architectures.

Emily: Welcome to The Business of Cloud Native. I'm Emily Omier, and I am talking with Dejan Deklich, from 8x8.

Dejan: So, I'm the Chief Product Officer at 8x8. To give you an idea, 8x8 is now 16 or 1700 employees worldwide, 450 million in revenue, give or take, offices all over the world, customers all over the world. I'm responsible for all product management, engineering, QA, project management operations for all the products worldwide for 8x8.

Emily: Can you give me a little bit of an idea of 8x8’s history in the Cloud?

Dejan: So, 8x8 has been around, probably, a lot longer than most companies you're talking about. We've been public 30 years, give or take. We have been in the business of communication and collaboration since early 2000s. As you can imagine, we have gone through so many different tech stacks, architectures, and so on, that it is pretty amazing.

We have, in the last several years, done a massive cleanup and rebuild of our software stack. We rebuilt pretty much all of the mobile apps, desktop apps, web apps. We rebuilt the platform starting with billing and provisioning all the way down to how the voice traverses the world. So, it's been a incredible couple of years, incredible journey where I would argue we have gone from the early versions of hosted service to early versions of Cloud, maybe 10 years ago, and we are now what I would like to call a proper cloud technology company. And it's been a very interesting, difficult journey. We learned a lot. We messed up a lot of things, then we learned some more than they did it correctly.

Emily: When you first moved to Kubernetes, and the modern public cloud, what was the rationale? What were their business reasons?

Dejan: Those multiple steps there. We moved to public cloud I don't know, five, six, seven years ago. We ran a lot of things in Amazon. And to be fair, we still also have data centers around the world. So, let me explain quickly what we actually running because I think it's important. So, we have, I think 16 data centers around the world, and then we run in pretty much every region of Amazon, we use Google Cloud extensively, and we have now shifted a lot of workloads to Oracle Cloud. At the same time, business is threatening me with Alibaba Cloud and Tencent Cloud as something that might be coming our way in the next couple of quarters. So, data centers are there because on the networking layer, the Cloud does not yet give us what we need for the realtime voice and video transmission.

We actually are the best voice provider in the industry. We have proven that, and that's where your milliseconds really matter, therefore networking still sits in data centers. As soon as the backbone can be moved into Amazon, and we are told that could happen in the next three to four years, we will move likely everything to the Cloud. So, what we have generally in the Cloud are different applications, and the reason for that is simply the velocity of deploying and scaling them.

So, what matters to us is, on one hand, the global reach: we have customers in 150 countries around the world. We have to have data centers close to the customers. And the applications need to be as close to the customer as possible, therefore all the different regions of Amazon, and Google, and whatnot. So, as you can imagine, managing all of that, monitoring all of that is a non-trivial exercise.

So, we moved to Kubernetes, in large reason, simply because it is one underlying framework that allows us to run workloads wherever we want. So, to give you an idea, we launched a video meetings product to compete with Zoom. We had, on launch, a couple of hundred thousand users, nothing really. And then, this COVID-19 happened, and within a period of weeks, we now hit 15 million users. The only way you can scale a system like that is if you have a properly built underlying architecture, everything horizontally scalable.

I was blown away, everything really worked. People were super busy, but by having proper cloud architecture, we were able to actually scale, and fulfill the demand that we have seen worldwide. Now, the nice thing is, as you put more and more workloads on top of Kubernetes, you can shift them between clouds as you want, or data centers as you want. And I think that's number one reason why we went with Kubernetes.

I love Amazon, I love Google, and nothing makes me happier than writing them a million-dollar checks, but I also want to be able to move the workloads wherever I can run them cheaply. And, to me, that's very important. I don't have unlimited budget; I have to be able to play the game and get the most compute and the most bandwidth for the lowest cost that I can, and Kubernetes lets me do that.

Emily: And would you say that Kubernetes was a technical decision or a business decision or both?

Dejan: That's a good question. I think normally, the way we operate at 8x8, you start with the business problem. The business problem was we don't want to be locked into one cloud. We want to be able to run wherever we want to run, and on top of that, we have customers in Europe who are not very friendly towards Amazon, and want us to run on other clouds. And then, we took a peek: what can we do? What's the fastest and easiest way to do it? Turned out it was Kubernetes, so that's the way we went.

Emily: What did the move to Kubernetes, what was it like? What were some of the surprises?

Dejan: It was very interesting. It is still very interesting. So, on one hand, the good thing was we have already broken the monoliths in the past God knows how many years, into services. But to get things running properly in Kubernetes, you have to go a bit deeper, you actually have to really clean up your code, and so on, and so on. So, one thing that I thought was incredibly useful was this allowed us to, for the first time in 8x8 history, create a proper template for a service where all your monitoring, logging, debugging, all of your stuff is standard for all the services.

So, as you can imagine, when you have a company that's been around for a long time, and you have a lot of services, and through various acquisition, you end up with different languages, different platforms, and so on, this now allowed us to clean up a lot of legacy and actually get to what I consider a really impressive engineering velocity. So, super challenging in many ways, because people had to say goodbye to some of the legacy tools, or ways of thinking, and so on. On the other hand, super-positive from the velocity we are now seeing. And honestly, if we haven't done some of these things, I doubt our platform would have survived the spike of the last two months.

Emily: What were some surprises in moving to Kubernetes? And you mentioned that you made some mistakes and learned a lot. What were some of those?

Dejan: It's funny, the good old story of automation and orchestration comes to bite you in the rear end every time, again. It's all nice when you have 10, 20 services. People can keep things in their mind. Once you are getting to hundreds or thousands of services that are running all over the world, things that you normally never want to do, such as documentation, such as fully automated CD/CI really become an absolute necessity. They stop being something that you learned about in your software engineering 101, and they become simply something that you really have to hardcore do.

And getting the engineering teams to accept that and changing the culture towards more documentation, more monitoring, more observability, was the key. It was also really interesting to see all these different monitoring tools, how they either couldn't scale, or if they were able to scale the cost would explode, or getting the insight out of the data that was being collected was next to impossible, or you needed somebody who spent their whole life looking into Splunk or something like that in order to figure out what's really going on. The release process had to be completely changed. We completely rebuilt—we blew up the CI/CD pipelines completely, rebuilt everything cleanly, and so on, and so on.

So, lots of very difficult engineering challenges, and I would argue the hardest thing was selling some of that work to the business. Business is always coming to you and saying, “Hey, we need feature XYZ, we need this product, we need that product.” And we would have to go back and say, “Well, actually, what we need is the availability, and the uptime, and stability, and we need to know what is really going on with the system so we can continue scaling it.” And looking back, I am really glad we did what we did because if we haven't started down this way, the 15 million users would not have happened.

Emily: Tell me a little bit more about making that sale to the business. Was there anything that's lost in translation, more about that conversation with the business about why this needed to happen?

Dejan: Right. So, it is very, very interesting. Most—and I have seen that pretty much in every company I've ever worked. There is always this feature bloat, and we are all guilty of that. And for sales, it is very easy to go to a prospect and go, “Hey, how do you like this?” Then prospects go, “Well, how about this and that feature?” sales says, “Absolutely.” And you end up with a roadmap which becomes this incredible sea-urchin-like monstrosity that goes in every possible dimension, and the cleanliness of the original product vision disappears over the years.

You end up with this extremely complex enterprise product, which does everything for everybody, and does everything badly for everybody. So, I don't believe that's the right way to do business. And I was lucky enough that my CEO and the board fully agreed with my view that we need to go back to the core; we need to figure out what's really important to our customers, and as bad as this COVID-19 is, and as much chaos as this caused around the world, it was really gratifying for us inside 8x8 to see that all the work we have done in the last few years was actually done correctly with a correct view of the future. Think about it, a few years ago, everybody was always debating, “Oh, should I have engineers working from home, or not? Are people more efficient working from home or not?”

Now, everybody's working from home. There is no more choice. There is no more talking about that topic. You have to be able to function from home. 8x8 builds the tools that ultimately allow enterprises to have their people function from home, both on the unified communication, and the contact center, and the video meeting side. So, you go, this now opens up the whole universe; I can hire people in anywhere around the world, we can do business anywhere around the world. And to me, that's the right way to go forward, but my job was to articulate what might happen, and what is likely to happen in the world and the economy in the coming five to 10 years, and this COVID-19 thing just pulled all of this work from home thing into now and not into next two, three, five years.

Emily: When you were talking to business stakeholders about moving to Kubernetes and why it was important. Was there anything that you felt like was lost in translation?

Dejan: They don't care about Kubernetes at all, and I never talked to them about Kubernetes. I talked to them about the scale and performance. Your CEO, your CFO, your Chairman of the Board, they don't care about Kubernetes or any technology. They care about the business problem. And the business problem is cost, velocity of innovation, global scale, global delivery, size of the engineering team, funding for R&D, you have to articulate what you want to do in the terms that matter to the business.

Technology is one of many ways to solve the problem. If you could have trained monkeys do something super fast and super cheap, it might be a valid way to solve a problem. As we are lacking trained, cheap monkeys, we have to use technology. And for a lot of us technology is a great, great toy that we love to play with, and we enjoyed dealing with it, so it is sort of a way we solve the problem. The way a business solves the problem is effectively through the P&L.

So, this translation between technical and business has to happen, and if you can't tell the story correctly, you will not get the funding you need, you will not get the project that you need, and ultimately the business will fail. And I think that's where the challenge lies. I see a lot of engineers and product people talk about the features, about how cool this or that would be, but they fail to forget that in the end, in a publicly-traded company, CFO and CEO go on the stage quarterly, and say, “This is the money we made. This is the money we spent.” And if you can help tell your story and the story of a product and engineering in the terms that they can use with the outside world, everybody's life becomes much easier.

Emily: How did that conversation progress as you progressed on the technical transition?

Dejan: So, it's like every other conversation. First time I said, “Yo, guys, we need to stop the feature work.” Obviously I got completely blank looks from everybody along the lines of, “Why would you ever want to do that?” after a couple of quarters of improving the engineering velocity, improving the product quality, stability, and so on, and so on. I think the whole company is now on the bandwagon. I think everybody by now understood why we are doing what we are doing.

Which is great. If initially, there was some pull back from Sales—Sales loves more features, it makes their life easier, but there is also, I have to tell you now, a very visible and noticeable pride in the whole company about what we were able to do, about the scale we have achieved, and how quickly and painlessly we actually got there. So, it was not easy, first few months, I have to tell you that.

Emily: What would you say were the biggest challenges?

Dejan: Changing the culture. Changing the culture for the whole company to stop chasing features and start chasing elegance, and speed, and stability, and uptime, and global scale, on one hand; on the engineering side, it was about, stop committing to featurettes, and start thinking about larger architectural projects that will massively improve the company; on the sales side, it was about how do I stop talking about the future roadmap and start talking about the vision of the company and the vision of the product. Those are very noticeable and hard to pull off changes, and it took a lot of effort from all the teams to actually get on the bandwagon.

Emily: Was there anything that was much easier than you expected?

Dejan: [laughs]. Probably the core technical parts. The nice thing with good engineers is that they always love to try the latest and greatest. So, I was worried that people here will have objections to changing some of the frameworks and so on. Turned out that that was absolutely not the problem. I was delighted that the engineers and PMs were able to jump on the latest and greatest and just go run with it, learn what they had to learn, then go forward. In previous companies, I have seen that there's much more of a challenge than at 8x8.

Emily: And then, from a technical perspective, what was more challenging than you expected, or what was a pain point?

Dejan: The CI/CD pipeline. It was really interesting how—it's almost once you have gone down the cloud way, you have to clean up how you build and release software. And most companies get to this sort of in-between state where there is a part of automation, but not everything is fully automated, and so on, and so on. You have to really talk to the people at really big companies to get to the point where there is proper really end-to-end CI/CD.

So, we were also in this in-between world. Yes, it's automated, yes it’s CI/CD, but not really. And we spent a tremendous amount of time and energy into cleaning that up and finding all the edge cases. And honestly, I don't think that work will ever end. It seems to be as soon as we finish one project, we find other things that can be improved and cleaned up, and then we go and do that. So, by far, the biggest challenge was the Continuous Integration. Getting that to the point where it really works is a lot, a lot of work.

Emily: Would you consider that among your biggest continuing challenges? Are there any other sort of challenges that you haven't quite figured out yet?

Dejan: I think this is one of the bigger ones because we are not the only AWS, or only Google Cloud, or only data center. We have a mix of everything, and a mix of different applications and platform components that have to reside all over the world. And it is stunning how much work goes into this. The other very interesting set of problems that we are encountering is on the data side, and figuring out—I would argue we haven't yet really fully figured out the data storage requirements.

So, to give you an idea—and it's not really so much even us figuring it out, it's the world figuring it out—every country has its own set of compliance regulations, security directives, they're all good, but how do you actually comply with all of them if you are providing service in 150 countries is a different problem. It almost feels like somebody on sand hill should invest in a startup which builds legal-as-a-service, something where the security and compliance will be done at the file system level so my engineers don't have to worry about it. To give you an idea, say you have a conversation, you and me: one of us sits in Europe, one of us sits in US. Who’s data privacy laws win? Where can the data reside? How do you store the data? What happens if you have an employee of a US company which resides in Europe? How is that data handled? So, there's a lot of these unknowns on the data security side, which are really interesting, which are really complicated. And I think if somebody manages to build a service around that, they will do very, very well.

Emily: Yeah. You have the employee of a US company who lives in Europe but is on vacation in Africa.

Dejan: Exactly. And you go, “Okay, does GDPR win? Does the US win? Does—” how do we do this?

Emily: Tell me just a little bit more about how Kubernetes and how the cloud-native architecture has made it possible to scale in the past month or so?

Dejan: Well, by removing all manual steps, and by having very clean containers that can be deployed anywhere, we were able to scale roughly 100x. And then, to make it even better, we were able to move from Amazon to Oracle Cloud. We did a very interesting collaboration with Oracle, where we agreed that we want to provide the highest possible security to our video product. We felt we can get that security and enterprise presence through Oracle Cloud, and not only did we scale up video product in the last I would say two months by 100x, we also moved very large parts over to Oracle in order to manage cost and get increased security. 100 percent possible through modern, latest and greatest architecture and automation. Otherwise, I don't know how we wouldn't have done it.

Emily: Has there been anything that surprised you over the course of this past two months, as you've seen the scaling. Any problems you encountered, for example, that you hadn't anticipated?

Dejan: It was interesting to see some of the monitoring tools have issues, as well as seeing the cost of all of these things really explode. It was also interesting to see the response of various cloud providers. As soon as we saw the cost explode, we reached out to everybody. It was really interesting to see how some are much more open to collaborating, others are less open to collaborating. It's been a very interesting two months.

But purely on the technical side, it's the monitoring that actually showed how important it is to be able to know what is really going on everywhere, and monitoring couple of thousand instances around the world, while possible is not necessarily easy.

Emily: In terms of things that went wrong, or things that went right, any surprises or any notable examples?

Dejan: Nothing—I was amazed that nothing really went wrong. I mean, we really worked our butts off, don't get me wrong. I mean, people have been up 16, 18 hour days for weeks to make sure everything works, and so on. But the fact that there were no serious problems still is amazing to me. I mean, 100x ramp in two months is a very nice ramp.

I was amazed with how well the team reacted to this. There was really no complaining. This is what we were all working towards: sort of the best possible technical time mixed with the worst possible personal time because of the virus. So, it's been really interesting to see how people all over the world really stepped up, and kicked butt, and helped us move forward.

Emily: Any comments you would have for other companies that are maybe a little behind 8x8 on the cloud transition?

Dejan: Yes, figure out your platform. If you don't have a good underlying platform for everything, you will end up paying for it. I'm hoping most people are done forklifting legacy products into Cloud. If you are still in that phase, really start thinking about building cloud-native. And it's not cheap, it's not easy, it's not trivial, but you have to do it if you want to survive. Otherwise, when something like this virus happens to you, you will not be able to scale, and you will fail at the worst possible time. Cloud-native is the way to go. It's the future. Embrace it, but don't think it will all be roses along the way.

Emily: Any other thoughts you have about the business of cloud-native, anything else you'd like to add?

Dejan: I think the very interesting thing on cloud-native is, cloud-native has many different forms, and thinking across multiple clouds, both public and private, I think is the way to go. I'm sure there are applications out there that are perfectly fine running in a single cloud and scaling in a single cloud. But I would imagine for the vast majority of enterprises, they are in a very similar world to 8x8, where you will be faced with challenges, across the whole world, which are better solved in different cloud providers at a different time. And the sooner you start thinking about your multi-cloud strategy, and how you will shift from one to the other, the better off you will be. And then, if you can throw in the private cloud into all of that, that's where I believe things become truly interesting, and probably the future is, in some form or the other.

Emily: All right. Well, I think we can go ahead and wrap it up there. Thank you so much for joining me. This was really great.

Announcer: Thank you for listening to The Business of Cloud Native podcast. Keep up with the latest on the podcast at thebusinessofcloudnative.com and subscribe on iTunes, Spotify, Google Podcasts, or wherever fine podcasts are distributed. We'll see you next time.

This has been HumblePod production. Stay humble.

View Details

Some of the highlights of the show include

  • The diplomacy that’s required between software engineers and management, and why influence is needed to move projects forward to completion.
  • Driving factors behind Ygrene’s Kubernetes migration, which included an infrastructure bottleneck, a need to streamline deployment, and a desire to leverage their internal team of cloud experts.
  • Management’s request to ship code faster, and why it was important to the organization.
  • How the company’s engineers responded to the request to ship code faster, and overcame disconnects with management.
  • How the team obtained executive buy-in for a Kubernetes migration.
  • Key cultural changes that were required to make the migration to Kubernetes successful.
  • How unexpected challenges forced the team to learn the “depths of Kubernetes,” and how it helped with root cause analysis.
  • Why the transition to Kubernetes was a success, enabling the team to ship code faster, deliver more value, secure more customers, and drive more revenue.

Links:

  • HerdX: https://www.herdx.com/
  • Ygrene: https://ygrene.com/
  • Austin Twitter: https://twitter.com/_austbot
  • Austin LinkedIn: https://www.linkedin.com/in/austbot/
  • Arnold’s book on publisher site: https://www.packtpub.com/cloud-networking/the-kubernetes-workshop
  • Arnold’s book on Amazon: https://www.amazon.com/Kubernetes-Workshop-Interactive-Approach-Learning/dp/1838820752/

Transcript
Announcer: Welcome to The Business of Cloud Native podcast where we explore how end users talk and think about the transition to Kubernetes and cloud-native architectures.

Emily: Welcome to The Business of Cloud Native. My name is Emily Omier, and I am here with Austin Adams and Zack Arnold, and we are here to talk about why companies go cloud-native.

Austin: So, I'm currently the CTO of a small Agrotech startup called HerdX. And that means I spend my days designing software, designing architecture for how distributed systems talk, and also leading teams of engineers to build proof-of-concepts and then production systems as they take over the projects that I've designed.

Emily: And then, what did you do at Ygrene?

Austin: I did the exact same thing, except for without the CTO title. And I also had other higher-level engineers working with me at Ygrene. So, we made a lot of technical decisions together. We all migrated to Kubernetes together, and Zack was a chief proponent of that, especially with the culture change. So, I focused on the designing software that teams of implementation engineers could take over and actually build out for the long run. And I think Zack really focused on—oh, I'll let Zack say what he focused on. [laughs].

Emily: Go for it, Zach.

Zach: Hello. I'm Zack. I also no longer work for Ygrene, although I have a lot of admiration and respect for the people who do. It was a fantastic company. So, Austin called me up a while back and asked me to think about participating in a DevOps engineering role at Ygrene. And he sort of said at the outset, we don't really know what it looks like, and we're pretty sure that we just created a position out of a culture, but would you be willing to embody it?

And up until this point, I'd had cloud experience, and I had had software engineering experience, but I didn't really spend a ton of time focused on the actual movement of software from developer’s laptops to production with as few hiccups, and as many tests, and as much safety as possible in between. So, I always told people the role felt like it was three parts. It was part IT automation expert, part software engineer, and then part diplomat. And the diplomacy was mostly in between people who are more operations focused. So, support engineers, project managers, and people who were on-call day in and day out, and being a go-between higher levels of management and software engineers themselves because there's this awkward, coordinated motion that has to really happen at a fine-grained level in order to get DevOps to really work at a company.

What I mean by that is, essentially, Dev and Ops seem to on the surface have opposing goals, the operation staff, it’s job is to maintain stability, and the development side’s job is to introduce change, which invariably introduces instability. So, that dichotomy means that being able to simultaneously satisfy both desires is really a goal of DevOps, but it's difficult to achieve at an organizational level without dealing with some pretty critical cultural components. So, what do I spend my day on? The answer to that question is, yes. It really depends on the day. Sometimes it's cloud engineers. Sometimes it's QA folks, sometimes it's management. Sometimes I'm heads-down writing software for integrations in between tools. And every now and again, I get to contribute to open-source. So, a lot of different actual daily tasks take place in my position.

Emily: Tell me a little bit more about this diplomacy between software engineers and management.

Zach: [laughs]. Well, I'm not sure who's going to be listening in this amazing audience of ours, but I assume, because people are human, that they have capital O-pinions about how things should work, especially as it pertains to either software development lifecycle, the ITIL process of introducing change into a datacenter, into a cloud environment, compliance, security. There's lots of, I'll call them thought frameworks that have a very narrow focus on how we should be doing something with respect to software. So, diplomacy is the—well, I guess in true statecraft, it's being able to work in between countries. But in this particular case, diplomacy is using relational equity or influence, to be able to have every group achieve a common and shared purpose.

At the end of the day, in most companies the goal is actually to be able to produce a product that people would want to pay for, and we can do so as quickly and as efficiently as possible. To do that, though, it again requires a lot of people with differing goals to work together towards that shared purpose. So, the diplomacy looks like, aside from just having way too many meetings, it actually looks like being able to communicate other thought frameworks to different stakeholders and being able to synthesize all of the different narrow-focused frameworks into a common shared, overarching process. So, I'll give you a concrete example because it feels like I just spewed a bunch of buzzwords. A concrete example would be, let's say in the common feature that's being delivered for ABC Company, for this feature it requires X number of hours of software development; X number of hours of testing; X number of hours of preparing, either capacity planning, or fleet size recommendations, or some form of operational pre-work; and then the actual deployment, and running, and monitoring. So, in the company that I currently work for, we just described roughly 20 different teams that would have to work together in order to achieve the delivery of this feature as rapidly as possible.

So, the process of DevOps and the diplomacy of DevOps, for me looks like—aside from trying to automate as much as humanly possible and to provide what I call interface guarantees, which are basically shared agreements of functionality between two teams. So, the way that the developers will speak to the QA engineers is through Git. They develop new software, and they push it into shared code repositories, the way that the QA engineers will speak to people who are going to be handling the deployments—or at management in this particular case—is going to be through a well-formatted XML test file. So, providing automation around those particular interfaces and then ensuring that everyone's shared goals are met at the particular period of time where they're going to be invoked over the course of the delivery of that feature, is the “subtle art,”—air quotes, you can't see but—to me of DevOps diplomacy. That kind of help?

Emily: Yeah, absolutely. Let's take, actually, just a little bit of a step back. Can you talk about what some of the business goals were behind moving to Kubernetes for Ygrene? Who was the champion of this move? Was it business stakeholders saying, “Hey, we really need this to change,” or engineering going to business stakeholders? Who needed a change.

I believe that the desire for Kubernetes came from a bottleneck of infrastructure. Not so much around performance, such as the applications weren't performing due to scale. We had projected scale that we were coming to where it would cause a problem potentially, but it was also in the ease of deployment. It had a very operations mindset as Zack was saying, our infrastructure was almost entirely managed—of the core applications set—by outsourcing. And so, we depended on them to innovate, we depended on them to spin up new environments and services.

But we also have this internal competing team that always had this cloud background. And so, what we were trying to do was lessen the time between idea to deployment by utilizing platforms that were more scalable, more flexible, and all the things that Docker gives with the Dev/Prod Parity, the ease of packaging your environment together so that small team can ship an entire application. And so, I think our main goal with that was to take that team that already had a lot of cloud experience, and give them more power to drive the innovation and not be bottlenecked just by what the outsourcing team could do. Which, by the way, just for the record, the outsourcing team was an amazing team, but they didn't have the Kubernetes or cloud experience, either.

So, in terms of a hero or champion of it, it just started as an idea between me and the new CTO, or CIO that came in, talking about how can we ship code faster? So, one of the things that happened in my career was the desire for a rapid response team which, that sounds like a buzzword or something, but it was this idea that Ygrene was shipping software fairly slow, and we wanted to get faster. So, really the CIO, and one of the development managers, they were the really big champions of, “Hey, let's deliver value to the business faster.” And they had the experience to ask their engineers how to make that happen, and then trust Zack and I through this process of delivering Kubernetes, and Istio, and container security, and all these different things that eventually got implemented.

Emily: Why do you think shipping code faster matters?

Austin: I think, for this company, why it mattered was the PACE financing industry is relatively new. And while financing has some old established patterns, I feel like there's still always room for innovation. If you hear the early days of the Bridgewater Financial Hedge Fund, they were a source of innovation and they used technology to deliver new types of assets and things like that. And so, our team at Ygrene was excellent because they wanted to try new things. They wanted to try new patterns of PACE financing, or ways of getting in front of the customer, or connections with different analytics so they could understand their customer better.

So, it was important to be able to try things, experiment to see what was going to be successful. To get things out into the real world to know, okay, this is actually going to work, or no, this isn't going to work. And then, also, one of the things within financing is—especially newer financing—is there's a lot of speed bumps along the way. Compliance laws can come into effect, as well as working with cities and governments that have specialized rules and specialized things that they need—because everyone's an expert when it comes to legislation, apparently—they decide that they need X, and they give us a time when we have to get it done. And so, we actually have another customer out there, which is the legislative bodies. So, they have to get the software—their features that are needed within the financing system out by certain dates, or we’re no longer eligible to operate in those counties. So, one of it was a core business risk, so we needed to be able to deliver faster. The other was how can we grow the business?

Emily: Zach, this might be a question for you. Was there anything that was lost in translation as you were explaining what engineering was going to do in order to meet this goal of shipping code faster, of being more agile, when you were talking to C level management? How did they understand, and did anything get lost in translation?

Zach: One of the largest disconnects, both on a technical and from a high level speaking to management issue I had was explaining how we were no longer going to be managing application servers as though they were pets. When you come from an on-premise setup, and you've got your VMware ESXi, and you're managing virtual machines, the most important thing that you have is backups because you want to keep those machines exactly as they are, and you install new software on those machines. When Kubernetes says, I'm going to put your pods wherever they fit on the cluster, assuming it conforms with the scheduling pattern, and if a node dies, it's totally fine, I'm going to spin a new one up for you, and move pods around and ensure that the application is exactly as you had stated—as in, it’s in its desired state—that kind of thinking from switching from infrastructure as pets to infrastructure as cattle, is difficult to explain to people who have spent their careers in building and maintaining datacenters. And I think a lot—well, it's not guaranteed that this is across the board, but if you want to talk about a generational divide, people that usually occupy the C level office chairs are familiar with—in their heyday of their career—a datacenter-based setup. In a cloud-based consumption model where it really doesn't matter—I can just spin up anything anywhere—when you talk about moving from reasoning about your application as the servers it comprises and instead talking about your application as the workload it comprises, it becomes a place where you have to really, really concretely explain to people exactly how it's going to work that the entire earth will not come crashing down if you lose a server, or if you lose a pod, or if a container hiccups and gets restarted by Kubernetes on that node. I think that was the real key one.

And the reason why that actually became incredibly beneficial for us is because once we actually had that executive buy-off when it came to, while I still may not understand, I trust that you know what you're doing and that this infrastructure really is replaceable, it allowed us to get a little bit more aggressive with how we managed our resources. So, now using Horizontal Pod Autoscaling, using the Kubernetes Cluster Autoscaler, and leveraging Amazon EC2 Spot Fleets, we were only ever paying for the exact amount of infrastructure that was required to run our business. And I think that is usually the thing that translates the best to management and non-technical leadership. Because when it comes down to if I'm aware that using this tool, and using a cloud-native approach to running my application, I am only ever going to be paying for the computational resource that I need in that exact minute to run my business, then the budget discussions become a lot easier, because everyone is aware that this is your exact run-rate when it comes to technology. Does that make sense?

Emily: Absolutely. How important was having that executive buy-in? My understanding is that a lot of companies, they think that they're going to get all these savings from Kubernetes, and it doesn't always materialize. So, I'm just curious, it sounds like it really did for Ygrene.

Zach: There was two things that really worked well for us when this transformation was taking place. The first was, Ygrene was still growing, so if the budget grew alongside of the growth of the company, nobody noticed. So, that was one really incredible thing that happened that, I think, now having had different positions in the industry, I don't know if I appreciated that enough because if you're attempting to make a cost-neutral migration to the Cloud, or to adopt cloud-native management principles, you're going to probably move too little, too late. And when that happens, you run the risk of really doing a poor job of adopting cloud-native, and then scrapping that project, because it never materialized the benefit, as you just described, that some people didn't experience. And the other benefit that we had, I think was the fact that because there were enough incredibly senior technical people—and again, I learned everything from these people—working with us, and because we were all, for the most part, on the same page when it came to this migration, it was easy to have a unified front with our management because every engineer saw the value of this new way of running our infrastructure and running our application.

In one non—and this obviously helps with our engineers—one non-monetary benefit that helped really get the buy-in was the fact that, with Kubernetes, our on-call SEV-1 pages went down, I want to say, by over 40 percent which was insane because Kubernetes was automatically intervening in the case where servers went down. JVMs run out of memory, exceptions cause strange things, but a simple restart usually fixes the vast majority of them. Well, now Kubernetes was doing this and we didn't need to wake somebody up in order to keep the machine running.

Emily: From when you started this transition to when you, I should say, when you probably left the company, but what were some of the surprises, either surprises for you, or surprises for other people in the organization?

Austin: The initial surprise was the yes that we got. So, initially I pitched it and started talking about it, and then the culture started changing to where we realized we really needed to change, and bringing Zack on and then getting the yes from management was the initial surprise. And—

Emily: Why was that a surprise?

Austin: It was just surprising because, when you work as an engineer—I mean, none of us were C suite, or Dev managers, or anything. We were just highly respected engineers working in the HQ. So, it was just a surprise that what we felt was a semi-crazy idea at the time—because Kubernetes was a little bit earlier. I mean, EKS wasn't even a thing from Amazon. We ran our Kubernetes clusters from the hip, which is using kops, which is—kops is a great tool, but obviously it wasn't managed. It was managed by us, mainly by Zach and his team, to be honest.

So, that was a surprise that they would trust a billion-dollar financing engine to run on the proposal of two engineers. And then, the next ones were just how much the single-server, vertical scaling, and depending on running on the same server was into our applications. So, as we started to look at the core applications and moving them into a containerized environment, but also into an environment that can be spun up and spun down, looking at the assumptions the application was making around being on the same server; having specific IP addresses, or hostnames; and things like that, where we had to take those assumptions out and make things more flexible. So, we had to remove some stateful assumptions in the applications, that was a surprise. We also had to enforce more of the idea of idempotency, especially when introducing Istio, and [00:21:44 retryable] connections and retryable logic around circuit breaking and service-to-service communication. So, some of those were the bigger surprises, is the paradigm shift between, “Okay, we've got this service that's always going to run on the same machine, and it's always going to have local access to its files,” to, “Now we're on a pod that's got a volume mounted, and there's 50 of them.” And it's just different. So, that was a big—[laughs], that was a big surprise for us.

Emily: Was there anything that you'd call a pleasant surprise? Things that went well that you anticipated to be really difficult?

Zach: Oh, my gosh, yes. When you read through Kubernetes for the first time, you tend to have this—especially if somebody else told you, “Hey, we're going to do this,” this sinking feeling of, “Oh my god, I don't even know nothing,” because it's so immense in its complexity. It requires a retooling of how you think, but there have been lots of open-source community efforts to improve the cluster lifecycle management of Kubernetes, and one such project that really helped us get going—do you remember this Austin?—was kops.

Austin: Yep. Yep, kops is great.

Zach: I want to say Justin Santa Barbara was the original creator of that project, and it's still open source, and I think he still maintains it. But to have a production-ready, and we really mean production-ready: it was private, everything was isolated, the CNI was provisioned correctly, everything was in the right place, to have a fully production-ready Kubernetes cluster ready to go within a few hours of us being able to learn about this tool in AWS was huge because then we could start to focus on what we didn't even understand inside of the cluster. Because there were lots of—Kubernetes is—there's two sides of it, and both of them are confusing. There's the infrastructure that participates in the cluster, and there's the actual components inside of the cluster which get orchestrated to make your application possible. So, not having to initially focus on the infrastructure that made up the cluster, so we could just figure out the difference between our butt and the hole in the ground, when it came to our application inside of Kubernetes was immensely helpful to us. I mean, there are a lot of tools these days that do that now: GKE, EKS, AKS, but we got into Kubernetes right after it went GA, and this was huge to help with that.

Emily: Can you tell me also a little bit about the cultural changes that had to happen? And what were these cultural changes, and then how did it go?

Zach: As Austin said, the notion of—I think a lot—and I don't want to offer this as a sweeping statement—but I think the vast majority of the engineers that we had in Seattle, in San Jose, and in Petaluma where the company was headquartered, I think, even if they didn't understand what the word idempotent meant, they understood more or less how that was going to work. The larger challenge for us was actually in helping our contractors, who actually made up the vast majority of our labor force towards the end of my tenure there, how a lot of these principles worked in software. So, take a perfect example: part of the application is written in Ruby on Rails, and in Ruby on Rails, there's a concept of one-off tasks called rake tasks. When you are running a single server, and you're sending lots of emails that have attachments, those attachments have to be on the file system. And this is the phrase I always said to people, as we refactor the code together, I repeated the statement, “You have to pretend this request is going to start on one server and finish on a different one, and you don't know what either of them are, ahead of time.”

And I think using just that simple nugget really helped, culturally, start to reshape this skill of people because when you can't use or depend on something like the file system, or you can't depend on that I'm still on the same server, you begin to break your task into components, and you begin to store those components in either a central database or a central file system like Amazon S3. And adopting those parts of, I would call, cloud-native engineering were critical to the cultural adoption of this tool. I think the other thing was, obviously, lots of training had to take place. And I think a lot of operational handoff had to take place. I remember for, basically, a fairly long stretch of time, I was on-call along with whoever was also on-call because I had the vast majority of the operational knowledge of Kubernetes for that particular team. So, I think there was a good bit of rescaling and mindset shift from the technical side of being able to adopt a cloud-native approach to software building. Does that make sense?

Emily: Absolutely. What do you think actually were some of the biggest challenges or the biggest pain points?

Zach: So, challenges of cultural shift, or challenges of specifically Kubernetes adoption?

Emily: I was thinking challenges of Kubernetes adoption, but I'm also curious about the cultural shift if that's one of the biggest pain points.

Zach: It really was for us. I think—because now it wouldn’t—if you wanted to take out Kubernetes and replace it with Nomad there? All of the engineers would know what you're talking about. It wouldn't take but whatever the amount of time it would to migrate your Kubernetes manifests to Nomad HCL files. So, I do think the rescaling and the mindset shift, culturally speaking, was probably the thing that helped solidify it from an engineering level. But Kubernetes adoption—or at least problems in Kubernetes adoption, there was a lot of migration horror stories that we encountered.

A lot of cluster instability in earlier versions of Kubernetes prevented any form of smooth upgrades. I had to leave—it was with my brother’s—it was his wedding, what was it—oh, rehearsal dinner, that’s what it was. I had to leave his rehearsal dinner because the production cluster for Ygrene went down, and we needed to get it back up. So, lots of funny stories like that. Or Nordstrom did a really fantastic talk on this in KubeCon in Austin in 2017. But the [00:28:57 unintelligible] split-brain problem where suddenly the consensus in between all of the Kubernetes master nodes began to fail for one reason or another. And because they were serving incorrect information to the controller managers, then the controller managers were acting on incorrect information and causing the schedulers to do really crazy things, like delete entire deployments, or move pods, or kill nodes, or lots of interesting things. I think we unnecessarily bit off a little bit too much when it came to trying to do tricky stuff when it came to infrastructure. We introduced a good bit of instability when it came to Amazon EC2 Spot that I think, all things considered, I would have revised the decision on that. Because we faced a lot of node instability, which translated into application instability, which would cause really, really interesting edge cases to show up basically only in production.

Austin: One of the more notable ones—and I think this is the symptom of one of the larger challenges was during testing, one of our project managers that also helped out in the testing side—technical project managers—which we nicknamed the Edge Case Factory, because she was just, anointed, or somehow had this superpower to find the most interesting edge cases, and things that never went wrong for anyone else always went wrong for her, and it really helped us build more robust software for sure, but there's some people out there with mutant powers to catch bugs, and she was one of them. We had two clusters, we had lower environment clusters, and then we had production cluster. The production cluster hosted two namespaces: the staging namespace, which is supposed to be an exact copy of production; and then the production namespace, so that you can smoke-test legitimate production resources, and blah blah blah. So, one time, we started to get some calls that, all of a sudden, people were getting the staging environment underneath the production URL.

Zach: Yeah.

Austin: And we were like, “Uh… excuse me?” It comes down to—we eventually figured it out. It was something within the networking layer. But it was this thing, as we rolled along, the deeper understanding of, okay, how does this—to use a term that Zack Arnold coined—this benevolent botnet, how does this thing even work, at the most fundamental and most detailed levels? And so, as problems and issues would occur, pre-production or even in production, we had to really learn the depths of Kubernetes. And I think the reason we had to learn it at that stage was because of how new Kubernetes was, all things considered. But I think now with a lot more of the managed systems, I would say it's not necessary, but it's definitely helpful to really know how Kubernetes works down in the depths. So, that was one of the big challenges was, to put it succinctly, when an issue comes up, knowing really what's going on under the hood, really, really helped us as we discovered and learned things about Kubernetes.

Zach: And what you're saying, Austin, was really illuminated by the fact that the telemetry that we had in production was not sufficient, in our minds, at least until very recently, to be able to adequately capture all the data necessary to accurately do root cause analyses on particular issues. In early days, there was far too much root cause analysis by, “It was probably this,” and then we moved on. Now having actually taken the time to instrument tracing, to instrument metrics, to instrument logs with correlation, we used, eventually, Datadog, but working our way through the various telemetry tools to achieve this, we really struggled being able to give accurate information to stakeholders about what was really going wrong in production. And I think Austin was probably the first person in the headquarters side of the company—I’m not entirely certain about some of our satellite dev offices—but to really champion a data-driven way of actually running software. Which, it seems trivial now because obviously that's how a lot of these tools work out of the box. But for us, it was really like, “Oh, I guess we really do need to think about the HTTP error rate.” [laughs].

Emily: So, taking another step back here, do you think that Ygrene got everything that it expected, or that it wanted out of moving to Kubernetes?

Austin: I think we're obviously playing up some of the challenges that we had because it was our day-to-day, but I do believe that trust in the dev team grew, we were able to deploy code during the day, which we could have done that in the beginning, even with vertically scaled infrastructure, we would have done it with downtime, but it really was that as we started to show that Kubernetes and these cloud-native tools like Fluentd, Prometheus, Istio, and other things like that when you set them up properly, they do take a lot of the risk out. It added trust in the development team. It gave more responsibility to the developers to manage their own code in production, which is the DevOps culture, the DevOps mindset. And I think in the end, we were able to ship code faster, we were able to deliver more value, we were able to go into new jurisdictions and markets quicker, to get more customers, and to ultimately increase the amount of revenue that Ygrene had. So, it built a bridge between the data science side of things, the development side of things, the project management side of things, and the compliance side of things.

So, I definitely think they got a lot out of trusting us with this migration. I think that were we to continue, probably Zack and I even to this day, we would have been able to implement more, and more, and more. Obviously, I left the company, Zach left the company to pursue other opportunities, but I do believe we left them in a good spot to take this ecosystem that was put in place and run with it. To continue to innovate and do experiments to get more business.

Zach: Emily, I'd characterize it with an anecdote. After our Chief Information Officer left the company, our Chief Operating Officer actually took over the management of the Technology Group, and aside from basically giving dev management carte blanche authority to do as they needed to, I think there was so much trust there that we didn't have at the beginning of our journey with technology and Ygrene. And it was characterized in, we had monthly calls with all of the regional account managers, which are basically our out-of-office sales staff. And generally, the project managers from our group would have to sit in those meetings and hear just about how terrible our technology was relative to the competition, either lacking in features, lacking in stability, lacking in design quality, lacking in user interface design, or way overdoing the amount of compliance we had to have.

And towards the end of my tenure, those complaints dropped to zero, which I think was really a testament to the fact that we were running things stably, the amount of on-call pages went down tremendously, the amount of user-impacting production outages was dramatically reduced, and I think the overall quality of software increased with every release. And to be able to say that, as a finance company, we were able to deploy 10 times during the day if we needed to, and not because it was an emergency, but because it was genuinely a value-added feature for customers. I think that that really demonstrated that we reached a level of success adopting Kubernetes and cloud-native, that really helped our business win. And we positioned them, basically, now to make experiments that they thought would work from a business sense we implement the technology behind it, and then we find out whether or not we were right.

Emily: Let's go ahead and wrap up. We're nearing the top of the hour, but just two questions for both of you. One is, where could listeners find you or connect with you? And the second one is, do you have a can’t-live-without engineering tool?

Austin: Yeah, so I'll go first. Listeners can find me on Twitter @_austbot, or on LinkedIn. Those are really the only tools I use. And I can't really live without Prometheus and Grafana. I really love being able to see everything that's happening in my applications. I love instrumentation. I'm very data-driven on what's happening inside. So, obviously Kubernetes is there, but it's almost become that Kubernetes is the Cloud. I don't even think about it anymore. It's these other tools that help us monitor and create active monitoring paradigms in our application so we can deploy fast, and know if we broke something.

Zach: And if you want to stay in contact with me, I would recommend not using Twitter, I lost my password and I'm not entirely certain how to get it back. I don't have a blue checkmark, so I can't talk to Twitter about that. I probably am on LinkedIn… you know what, you can find me in my house. I'm currently working. The engineering tool that I really can't live without, I think my IDE. I use IntelliJ by JetBrains, and—

Austin: Yeah, it’s good stuff.

Zach: —I think I wouldn't be able to program without it. I fear for my next coding interview because I'll be pretending that there's type ahead completion in a Google Doc, and it just won't work. So, yeah, I think that would be the tool I'd keep forever.

Austin: And if any of Zach's managers are listening, he's not planning on doing any coding interviews anytime soon.

Zach: [laughs]. Yes, obviously.

Emily: Well, thank you so much.

Zach: Emily Omier, thank you so much for your time.

Austin: Right, thanks.

Austin: And don't forget Zack is an author. He and his team worked very hard on that book.

Emily: Zack, do you want to give a plug to your book?

Zach: Oh, yeah. Some really intelligent people that, for some reason, dragged me along, worked on a book. Basically it started as an introduction to Kubernetes, and it turned into a Master's Course on Kubernetes. It's from Packt Publishing and yeah, you can find it there, amazon.com or steal it on the internet. If you're looking to get started with Kubernetes I cannot recommend the team that worked on this book enough. It was a real honor to be able to work with people I consider to be heavyweights in the industry. It was really fun.

Emily: Thank you so much.

Announcer: Thank you for listening to The Business of Cloud Native podcast. Keep up with the latest on the podcast at thebusinessofcloudnative.com and subscribe on iTunes, Spotify, Google Podcasts, or wherever fine podcasts are distributed. We'll see you next time.

This has been HumblePod production. Stay humble.

View Details

Some highlights of the show include

  • The challenges of operating digital commerce at scale, including the need for resource pooling and resiliency — and how this caused Ant Financial to re-think their infrastructure.
  • Ant Financial’s former approach to scaling, which was mostly manual, and highly resource-intensive.
  • How Kubernetes is expediting cloud development for Ant Financial.
  • Haojie’s thoughts on the global engineering skills gap, and China’s growing cloud computing market including driving factors and barriers.
  • Why Ant Financial’s migration has largely been a success — and why achieving operational security is now a top priority for the company.
  • How Ant Financial is managing disconnect between its engineers and business leaders.
  • The company’s ongoing mission to migrate its systems and applications away from legacy architectures.

Links

  • LinkedIn: https://www.linkedin.com/in/haojiehang/
  • https://www.investopedia.com/tech/worlds-top-10-fintech-companies-baba/

Transcript

Announcer: Welcome to The Business of Cloud Native podcast where we explore how end users talk and think about the transition to Kubernetes and cloud-native architectures.

Emily: So, I always start the same way. Can you introduce yourself?

Haojie: Hey, my name is Haojie Hang. I'm a product manager in the CTO office at Ant Financial. I work on the product and strategy side for, basically, the CTO and the other executive leaders, as well as leading a small product teams within the org to look at the frontier technology in the cloud and other infrastructure businesses.

Emily: And can you tell me a little bit more about what Ant Financial does? And then, also, what do you do on a day to day basis? What do you do when you get into the office?

Haojie: Yeah, I'll do a quick introduction about the Ant Financial business. It's not just one business or two business, it's a group of businesses that we innovate and we do, mostly in China, but we're also expanding very rapidly all over the world. So, Ant Financial is basically a group of businesses including credit for both consumers and the enterprise, as well as loan businesses, both consumer and enterprise businesses. We say that the parent organization is basically, we call it Alipay, it’s the earliest business we do since 2004 when the business was basically born from Taobao, which is our parent company. So, in short, the Ant Financial Business has a lot of presence in the business of payments business, remittance, credit card, loans, securities, and many other businesses like intelligent technology, blockchain, pretty much everything you can imagine in the FinTech and financial services, we’re in there.

Emily: Tell me a little bit more about the cloud-native journey for Ant Financial. When did it start? Why did it start? What was some of the motivations behind moving to cloud-native?

Haojie: Yeah, it's actually quite interesting. I joined Ant Financial in 2008, but actually, the entire company started to look at cloud-native technology quite early, in 2012. So, back then, people were just looking at these technologies around the world, mostly from the US, they look at this open-source community, look at what other companies are doing, how to use the cloud-native technology to help with their business in the peak time, so during event. There’s online promotion event we're doing every year, called Double 11—Shuāng shíyī in Chinese. Every year, so we have a large amount of promotional events happening online, trying to help merchants and the customer is trying to sell and buy stuff in our Tmall and Taobao platform in very, very discounted price. So, for that promotion event online, we have to think about the resilience, the resource pooling, oftentimes the visits has to increase multiple times, sometimes over 100 times the increase compared to the normal time. So in that case, we have to think about how we can be very resilient and efficient infrastructure to support that business needs. So, this is a very large topic. And then, back then, there was a lot of focus and study in our cloud computing department. So, we started looking at this technology called Mesos in 2012. And then, we do a lot of experiments around this technology, but from the business perspective, it's still hard to justify the benefits of moving to Mesos completely. So, we have multiple teams doing a lot of research in Mesos, in Kubernetes, sometimes in our own technology stack, but there's not enough proof or enough confidence for us to move completely over to that technology, until the emergence of Docker container, this Docker technology. Then we started to look at our container infrastructure, really do the investigation around this technology, and understand why this is taking over so quickly over the world, from the business perspective, and from the technology perspective. If you look at the community of Docker, the thing does not really happen until 2015. But we are already in the game for about a year or two. So, we're actually quite happy about our original strategy, but it's just in terms of the research. We're actually a little bit behind in terms of moving to this cloud-native architecture. But as you can see, that I had an interview with CNCF. So, we are very happy about the results that we have right now. Pretty much the entire architecture we run within Ant Financial is, basically, on Kubernetes ecosystem. It's not just using the open-source version of it. We're doing a lot of customization around this open-source framework. Yeah, I can talk more about the details.

Emily: Yeah. Well, let's back up just a little bit. I’m curious what you were doing to manage this scaling before? And how did that change? And what about the whole process changed? Like, how stressful is it now, compared to before?

Haojie: The process was very manual, I would say. We have extremely large team of engineers, and DevOps, security teams. And oftentimes their responsibility are overlap. So, some engineers are doing security work, some engineers are doing basically operational work. I would say, some people really hated it because they have to be on the computer, look at monitor 24/7, making sure transactions succeeded. When the peak time happens, there's nothing wrong with it. Sometimes they have to keep their phone open 24/7, basically to make sure this thing will not fail, right? And then, just many parts of work has to—so in the previous way, the way we do this operation is quite manual. We don't have a mature system or methodology telling us what we should do first, we should do second, and what's what would you do after this. So, basically the collaboration chain was not there. Therefore, when issue happens, our operation team has to respond very quickly. But then, how can we quickly identify the problem, and make it a problem? That's a problem, right? So, we have to make sure every time we respond, we respond in a very effective manner. That's the problem. In the previous process when something unexpected happen, who had to engage with the entire team from product, engineering, operation, security, everybody has to get up and look at the problem together, which was quite inefficient. So, after we moved to this cloud-native architecture—it's not the standard cloud architect, it's, kind of—we have a lot of innovation on top of this, to make sure that’s fitting to our tech community, to our businesses. So, we basically did a lot of innovation in the process to make sure after we had this transition, people are clear about the roles, what they should be responding, and then who should be doing what. That's quite important.

Emily: And tell me a little bit more about some of the additional layers or some tools that you've built on top of Kubernetes, and how that's helped you be successful.

Haojie: Yeah, so I can give you an example. So, when we look at Kubernetes technology in our intelligent technology, or intelligent cloud business units, we're thinking about how can we use Kubernetes for cloud deployment. Okay, so previously we are using Mesos to do that. But we found this technology that lots of people are familiar with this technology. And then, people are not very sure if Mesos is the right path for container management, resource management or cloud deployment. But when we move to the Kubernetes for cloud deployment, people are actually quite happy. We are seeing a decrease in the amount of—to stand up the cloud. And previously, it took us two to three weeks to build an entire cloud. But after we use the Kubernetes technology, we can do that in a week. Oftentimes, if the scale is smaller, we can do that within three days. That's quite important because people are confident about the community; about this technology. And then, from the users perspective, they are also more willing to invest in Kubernetes. Oftentimes, this is the chicken-egg problem, right? When more companies are hiring more this, these terms appears in the market, in the job descriptions, the more people are willing to learn. So, this is actually what we're seeing a very, very good cycle for me, from both a company perspective and the talent perspective. So, that's actually quite good. But the problem for us is, there's still not enough people, or we say, you know, good talents in the market that we can attract. Basically, we're seeing a shortage of great engineering talent in the market, after the cloud-native transition. So, we're still trying to think about how we can educate the internal audience in the technology community to help them quickly pick up this new technology in the cloud, as well as the practice behind the cloud-native architecture.

Emily: You know, I wanted to talk to you a little bit about the overall situation in China and that there's also this, sort of, skills gap. It sounds like it's just as present in China as it is in other parts of the world.

Haojie: So, I would say in terms of the cloud computing, cloud-native tech community, it's pretty much—we had community forming as early as the rest of the world. But then, in early days, it was just a marketing term by people saying, oh, this is cloud, we want to learn something. But then really, from the business perspective, there's still not enough customer trying to pay for this technology. Oftentimes, the contract size was not large enough to feed engineers. That's what I say. And then, I think the trend of the serious adoption really happens in the past three to four years when a lot of startups coming out in the cloud businesses. There’s a company called [QingCloud]. There’s a company called [unlcear]. There’s many other unicorns in China in the cloud computing space, and I think two of them just went public in the A-listed share recently. So, from that perspective, I would say, the cloud computing business is really maturing rapidly in the past two years. Because we see some unicorns really coming out of this game, besides Alibaba. I think, from that perspective, I will say, it's getting better and better. It's just in terms of the pay behavior, right? How much customers are willing to pay for this technology, pay for the services, pay for the products? I think it will still take some time to mature.

Emily: To what extent do you see Chinese companies using cloud-native services and tools from Europe, from the United States, from elsewhere in the world? To what extent is it a segregated market? Like, the rest of the world doesn't use Chinese tools, Chinese companies don't use tools from the rest of the market.

Haojie: It's a good question. To me, I'm coming from an engineering background, I believe open-source community is global. It's a global phenomenon. I think the world is connected. And I would never say that people in China, or Chinese companies, or—they only using technology businesses that created in China, this is not right. And often in many cases, there's not enough options, right? So, I think, even though Chinese companies and startups are trying to innovate very aggressively, but I think the world is still connected, they have to build on top of the innovation that’s already happened in the rest of the world. So, in that case, I think we're still seeing a lot of collaboration across the globe. China, United States, for sure, Europe, other parts of the world. It's just how aggressive people are in terms of investing in frontier technology. And are they really seeing the benefits of using the frontier technology? There’s question of technology innovation versus business innovation, right? Do you see the business value? Can you really see that in the next five years? I would say in China, most of the non-internet sector, they're quite short-sighted. They're still trying to survive, they're trying to make sure they are doing—they can become the top three, top five in their business. So, technology is oftentimes secondary. But for the leaders in that sector, they have to think about that quite early in order to become the top player in their sector. So, I think, the trend is that people are still collaborating academically and engineering side to make sure the right technology gets applied in the right scenario and trying to improve the technology at home.

Emily: It's interesting that you mentioned that some Chinese companies might not focus as much on the technology. Do Chinese companies tend to consider moving to cloud-native, important for their business? Or a strategic move?

Haojie: From the strategy perspective, yes. Every leader would definitely know about cloud computing, cloud-native architecture, they would definitely think about moving. It's just that internal execution when they think about moving seriously, they have to evaluate, do we have enough talent? How much business value am I getting out of this? Is it really helping? What is my budget? And all those kind of fears, problems. So, that's what backing them up because oftentimes they don't have enough budget. That's what I say in those non-internet sector. Because I've lived and worked in the US for a while. I think in the US, the non-internet sector are quite advanced in terms of the technology adoption, especially in cloud-native. It's quite easy for them to recruit, then build a large engineering team to work on cloud infrastructure software. But it's not the case in China because people are still trying to, especially the leader who can make the decision, they're still thinking about the ROI, the rate of returns, the rate of investments, for building a strong software teams, making sure they have the robust infrastructure running at the bottom level. So, they are still trying to figure out the budget, make sure they are profitable enough to afford that.

Emily: Do you feel like Ant Financial got the business benefits out of the transition that it was looking for?

Haojie: Yeah, as I mentioned, the entire organization are quite happy about the move because really, they are, kind of, [unintelligible] move in China. So, basically, even the non-engineering teams started to appreciate this, and talk about this technology, and trying to understand it deeper, because they see the entire organization are quite happy, especially from the business protective. As I mentioned, in the Double 11 event we have—last year in 2019, the GMV we had was 260 billion RMB in total, which is 25 percent growth compared to the last year. So, for that large amount of GMV, we supported the entire infrastructure, are building from our cloud infrastructure. It is quite massive. We don't have a infrastructure for that business, we have the infrastructure for the entire group, including Ant Financial and Alibaba. Basically the entire businesses is running in the cloud. We have very, very few siloed data centers and the infrastructure—you know, uh, data centers—basically, we have the entire thing running in single cloud. That's the largest achievement we had I think since 2019, which was one of the strategic goal we had, we achieved last year. And this year, we're putting a lot of emphasis in the secure operation. It has been one of the primary cloud business goal because when bad things happen, people are literally losing money. Imagine one of the transactions failed. It failed the entire country, right? Like, no one else in China or in the other part of the world can make a purchase from Alipay app. This is quite devastating. So, secure operation has been the only thing we focus on this year, I would say. I remember in some meetings, one of the leader mentioned, “If there’s only thing we should do this year, it’s secure operation.” We're trying to make sure we operate the entire business safely on top of our new cloud-native architecture, with the minimum amount of incidents and failures.

Emily: And what do you think have been some of the challenges? What has been more difficult than you imagined in making this transition?

Haojie: Yeah, I think for me, the most obvious point is that we still have a large amount of operating team engineers, and support team, and the product, and the entire organization, basically, to making sure the entire thing working seamlessly because I think it's very hard to quantify it. I think the overall efficiency in running the cloud-native architecture, we're still looking at that. Let me try to find a good example.

Emily: Let me ask a question. What's gone unexpectedly well? Was there anything that you thought was going to be really challenging that wasn't?

Haojie: Oh, I think after moving to the cloud-native architecture, the engineers are quite happy. They're working much, much harder. They're trying to do things much more quickly than we imagined. Basically, they are very aggressive, and very happy to see the leadership teams really buying this technology, and they’re invest—want to invest seriously in this technology. They are building not only the engineering team but also the prod team, the entire organization around to cloud-native technology. So, oftentimes in order to persuade business leaders to do something serious in the technology, they have to spend a lot of time trying to evangelize to the leadership team to making sure they understand, oh, this is the right direction. We have to do this right. It takes oftentimes from six months to a year for them to really doing that. So, for that, I think it's quite successful. We see a very—basically I think the entire engineering culture has changed. People are looking at open-source community more aggressively. They think about how we should contribute back to the community. What community events should I support? What conferences should I go to? There's more and more discussion like that happening within the organization. And, I think, larger Ant Financial has become one of the sponsor in the events. We are one of the most active participants in the community, I think, since 2019, along with Alibaba. So, that's the positive side I'm seeing. People really start to form a culture on their own, especially in open-source community. Trying to be more present, trying to take more active position in the discussion, both within company and outside of company. So, that's actually quite the good. We're happy to see engineers are doing their work, and are doing it more aggressively.

Emily: Do you feel like there's any sort of disconnect between the engineering teams and business leaders? Or do you feel like they're mostly on the same page?

Haojie: Yeah, I would say there are still some gaps between the business leaders and the engineers. So, oftentimes, I would say the engineers are quite updated with what's going on in this community, in some new plugins, in some new components coming out of this Kubernetes ecosystem, but then the business leaders don't have enough time to to pay attention to this. So, it really depends on how confident they are about this technology. And how much more time do you want to put into this personally. I think the business leader will look at the numbers like KPIs, metrics, the number of accidents, the operating efficiencies, things like that, but that’s in the business context, all right. The engineer leaders cares more about what kind of new technology we use, what kind of new technology we created on top of this ecosystem, and how many people are happy about using this technology? And how many more can we do from this transition? So, basically, they are disconnect. So, I think the good part in Ant Financial is that for business leaders, most of the business leaders are coming from engineering background, but they have a strong KPI in their work. And then, most of the engineer leaders has to learn business, because, in order to persuade business leaders to invest in this, they have to think from their perspective. I think, in terms of the communication, they're quite up to date. It's just in terms of the execution and the timelines there are some disconnected happen. Yeah.

Emily: What would you say that business leaders are looking for that engineering teams might not be thinking about?

Haojie: I think one example that I see is a business leader will think about the team building, the talent building, the culture, and the public image that we had in the public, especially in China. Yeah, let me give you an example. If the technology—if the company is not cool enough, from the technology, from engineering perspective, it’s very hard to attract the top talent in engineerings from the business leader. Without strong engineering teams, we cannot execute. We cannot innovate. So, that's something they oftentimes think about when they try to invest in technology. But in terms of the execution, after the engineers gets on board, and work in Alibaba, in Ant Financial, that's something engineers have think about. How do they keep the talent? How do they make sure talents are happy? How do we make sure they are satisfied about what they do? So, I would say these two things have to work at the same time. You cannot have a strong image in technology, in frontier technology. But then, after the talent gets on board, they realize, oh, this is just great from outside, but from inside, we are still working on the legacy technology. It's operate very inefficiently internally, and how can we make sure people are dealing this? And I think that's quite important.

Emily: Is there anything that you think is preventing you from moving further along in the cloud-native journey? Anything other than lack of human resources?

Haojie: I would say that how can we securely move away from the legacy architecture, whether it's built privately or you built it using other vendor’s technology? You know, for that kind of transition we're taking very seriously. We still have a large amount of systems and applications running on Oracle, running on, sometimes in MySQL, sometimes in other siloed stack. And we're not 100 percent. We're in one cloud, we're not 100 percent away from Oracle, MySQL or that type of, we consider legacy, architecture. So, the moving will still take some time. And so, how can we make sure the transition is successful? How can we make sure the transition is less painful? Is something we as the leaders and the business executives will think about because how do you how we can set up the right KPI and the right goal for engineers to feel happy about doing this work? I think that's one of the challenges. Oftentimes people, when they are placed into this kind of work, moving from legacy architecture, to new architecture, just very minimum business value we can see from this transition, right? So, we have to have the right—we have to set the enough goal to motivate them to do the work. That's something we have to really think about that in the long term. Because this is not like we do that for six months, a year. It's going to be an effort for the next three to four years. Imagine, Alipay business started in 2004, and it's been already 16 years. So, the transition was to happen over time. It's just, how we can make sure that the transition that it's less painful?

Emily: Tell me just a little bit more about some of the custom capabilities that you built on top of Kubernetes.

Haojie: We have our own internal monitoring architecture, which is quite advanced, I would say. And this kind of monitoring infrastructure is built for both developers and operators. I think that is something we invest extremely heavily because we cannot find any other alternatives in the market. I'll give you some background about this monitoring infrastructure. So, the entire tech stack was primarily built on Java stack, the thing starting from 2004. And now a lot of cloud-native technology are leveraging Go technology, right? So, the monitoring of Go is quite different from the monitoring of Java. We have different versions of JDK and JRE that we created—one of them was actually recently open-sourced called Dragon Well. You can check out on online, a lot of posts around that. So, we have to make sure the entire stack, from the application, middleware, in the mesh-level, container host, all the way down to compiler has to be monitored quite efficiently. Once anything happened, from the operation side or from the technology side, we have to quickly respond to identify in what layer the error happened. In order for that mitigation to be efficient, we have to make sure we are monitoring every single thing in the stack. As I mentioned, from the application, middleware, host level, all the way down to hardware level, sometimes a failure in hardware will cost the entire failure in our business. It's quite often. So, we have to make sure we are monitoring our own technology in a very good manner. And also imagine monitoring that amount of infrastructure in that massive scale. It's very challenging. I think before 2014, we had a lot of failure in our monitoring infrastructure. This is quite ridiculous, but this is what happened. So, we spent a lot of time to make sure we have the supporting infrastructure ready for that kind of businesses. That's quite important.

Emily: Anything else that you'd like to add about either your own experience moving to cloud-native or some observations about how things are going in China in general?

Haojie: I think from the strategy perspective, Chinese company or startups from China are doing quite well. It's just the market is quite different. For companies to survive and thrive in the Chinese market, they have to go with the customers, right? So, even though the innovation happens at the same level, the customers are not at the same level from what I see. But overall, I think the trend is quite positive, I think eventually, be it five years, or seven years, or ten years, Chinese companies, Chinese customers will be at the same level as the rest of the world: in the US, in the UK, in Australia, in the rest of the world. I think people are more and more aggressive, and they would like to allocate more and more budget into technology business. They realize the benefits of it, especially in the current outbreak. When people, they cannot go to work, but they still have to do something. The business has to survive. Like, they have to do something in order for the business to survive. So, from the business perspective, how can they build their strong online presence during the outbreak? Is actually quite important. Before the outbreak, I would say, in the retail business, there still some people think about, “Oh, how can we do this in our traditional manner? How can we open as many stores as possible.” They didn't really care about building a store online. From in Taobao or Tmall, [unintelligible] seriously. But during outbreak, people they have to stay at home. They have nowhere to go. But then the business, they still have to pay their employees. So, how can you do that? The only thing is going online. In order to go online, they have to build online infrastructure for their customers, for their employees, for them to work. So, that's quite—honestly, that's one of the trend I'm seeing: that people are paying more and more attention to work remotely, and use software, SAAS software without on-premise deployments. In that case, people, they are able to work wherever they go. Being at home, office, on the road, people are really interested in the benefits of SAAS, of cloud. I think that's something that I'm seeing. I think after this year, definitely the market of SAAS will become better and better because not only the technology is, but the business leaders will understand the value of using Zoom, using Ding Ding, using WeChat, to make sure their employees, they can work anywhere they want.

Emily: Well, thank you so much. A couple finishing up questions. First of all, what is an engineering tool that you couldn't do your job without?

Haojie: Do you mean, like, just tools for me to do some engineer work?

Emily: Yeah. What's your favorite tool, something you just can't imagine working without?

Haojie: We have a lot of tools innovated within the company. I don't think I can mention that in this podcast.

Emily: Okay. No problem. And then, how can people connect with you if they want to?

Haojie: At work, or outside of work?

Emily: like on Twitter or on social media.

Haojie: Yeah, I had a lot of invitation from LinkedIn, not so much on Twitter because I'm not active on Twitter. But I think people, they get to know me, oftentimes from word of mouth, they got introduced from other friends of mine, they want to understand about the technology adoption in China, especially in the cloud. Yeah, people oftentimes, which me from LinkedIn, that's the primary source.

Emily: Well, thank you so much. I really appreciate you taking the time to chat.

Haojie: Thank you, Emily.

Announcer: Thank you for listening to The Business of Cloud Native podcast. Keep up with the latest on the podcast at thebusinessofcloudnative.com and subscribe on iTunes, Spotify, Google Podcasts, or wherever fine podcasts are distributed. We'll see you next time.

This has been HumblePod production. Stay humble.

View Details

Some of the highlights of the show include

  • How containerization enabled Nav to spread roughly 250 virtual machines across multiple environments, while drastically reducing infrastructure spend
  • Travis’s thoughts on buying cloud native software tools versus building them, and what engineers should consider during this process
  • The difficulty of finding security solutions that work inside of a cloud-native ecosystem
  • Why companies should expect to encounter unique challenges when migrating to Kubernetes
  • Why companies need to understand their end goal, and determine an overall objective before beginning a migration
  • Travis’s must-have engineering tool, and why he can’t live without it

Links

  • LinkedIn: https://www.linkedin.com/in/stmpy/
  • Twitter: https://twitter.com/stmpy

Transcript

Announcer: Welcome to The Business of Cloud Native Podcast where we explore how end users talk and think about the transition to Kubernetes and cloud-native architectures.

Emily: Welcome to The Business of Cloud Native. I’m Emily Omier, your host. And today I’m here with Travis Jeppson. Travis is currently at Kasten, but he’s also going to talk about his time as a director of engineering at Nav.

Travis: At Nav, my role shifted quite a bit while I was there. I started as a software developer, writing Ruby back end applications for them, and then shifted into—actually within a month of being there, they shifted me over to the operational side because I had previous experience working with containerization, and also in infrastructure. So, they quickly moved me over into that realm and from there, I worked there for about a year until they told me, go spin up a team and get things moving. Help us move to containerization. Help us move to a more modern infrastructure and stuff. And so, about a year after that I became a director of engineering to where I had our ops team that had spun up, and then I also acquired both our QA team and our IT team that was there. And then, about a year after that, I ended up acquiring a little bit more than that. So, I ended up with a fair amount of our front end and some of our backend teams as well, and where they moved me into the senior director position. So, a day in the life, towards the end of when I was at Nav was a lot of working with the teams, helping them to do a lot of architectural perspective, and changes, and outlook to where we were trying to get as far as the company is concerned. We were building a product that we could address both first-party customers where they would log in to the Nav website directly, as well as working with partners so that we could issue out Nav functionality to those partners that they could incorporate to their pages as well. And so, we worked very hard to try to segment those two pieces together so that what we were building could be dispersed between both first-party customers and our third-party customers. And so, towards the end of my time there, it ended up being a lot of working within all of engineering to help facilitate those purposes. Then, just about six months ago, I ended up shifting my role over to a company called Kasten. And, Kasten is strictly working within the Kubernetes ecosystem. So, we do data management for Kubernetes based applications, and I am the site lead in Utah for Kasten, and so my day in and day out, a lot is, it's, kind of, all over the place. Sometimes it's working with engineering to help figure out some things going on there, sometimes it's working with brokers to help find office space for it. And sometimes it's dealing with insurance. It ended up being quite dynamic. But overall, I'd say most of my time is really spent more on the engineering side, just from the perspective of having worked at Nav and having been a consumer of a lot of these technologies, I think that they really appreciate my insights that I'm able to give there. So, I end up working, a lot, with the engineers to help facilitate what we're doing.

Emily: Sounds like you end up serving as a bridge from having been an end-user. But do you think that there is common miscommunications that happen, or what do those conversations sound like? Why is that experience valuable?

Travis: Yeah, so I don't know if it's as much as a miscommunication as much as what are customers looking for? And what are they trying to achieve? And why are they purchasing different software solutions? And what makes sense for them, more than anything. And I think that, having been a consumer of those products, I was more or less on the front lines there. When I was building our operational team at Nav, that was basically what I was doing is trying to figure out what things are we going to spend time on? And what things are we going to build ourselves, or what things do we need to just go find a solution for and bring them in-house? And the funny thing is when I was doing that for Nav is actually when I was introduced to Kasten and to the CEO here. And so, that ended up changing the way my career went. But overall, I think what Kasten—what those conversations really end up becoming is what are customers trying to do, and where are they trying to go?

Emily: Yeah, and in fact, that is exactly what I want to talk about more on this podcast. So, tell me a little bit about what your experience at Nav was. What were you looking for? What did you want to prioritize? What was the company hoping to get out of moving to containers?

Travis: So, I would say maybe the piece that really facilitated a lot of the progress in that sense was starting to understand our infrastructure spend. And then, to couple with that was also trying to become more agile. More agile in the sense of being able to push on demand, where previous to that we were pushing—you know, when we push our code, we did it on a bi-weekly basis—well, every other week, and it was always very cumbersome. If we have pictures of us in the early days of Nav, where there would be 10 engineers around someone’s desk, and they were the one person that was pushing the code into production, just waiting for the other shoe to fall, or waiting for something to happen. And so, when I started doing operational things for Nav, it started addressing those two things. What can we do to help control our infrastructure, and to understand it a little bit better? And how can we also create more of a dynamic infrastructure? Like, Nav is very much a US-based company. And so, the traffic that we're getting onto our website was regional very, very much. And so, there would be periods where it would be very busy, and then there'd be periods where it wasn’t. And the way that our infrastructure was designed, and a lot of times the way that they are designed, especially with virtual machines, is that you're building for capacity. You're building to be able to handle that load, and that has to stay there all the time, regardless of whether that capacity is being used or not. And so, that was one of the biggest questions, and that bill was—we were completely in the clouds. We were completely in AWS, but that bill continued to get more and more expensive every month. To the point of where it warranted the executive team to come down and say, “This needs to be fixed. This is going at an outrageous pace, and we need to be able to figure out how to control this.” And so, that's when they came to me and said, “Okay, get a team spun up, and let's figure out how to control this.” And so, I would say that those were some of the big pieces that really drove us to start looking at cloud-native technologies, containerization, and Kubernetes.

Emily: And do you think it was successful?

Travis: Yeah. So, I do, for a few reasons. And obviously, we learned some lessons along the way, but what we were able to do is, with the infrastructure that we had growing, we were pushing close to 250 virtual machines across two different environments, that being our production, and a development environment that we had. And when we moved to containerization, we were able to not only spin up more environments, but we were able to still decrease that overall spend as far as the infrastructure was concerned. And so, what used to be, I think we had about 100 VMs for our dev environment and then about 150 for our production environment, and that crossed many different pieces from the front end to the back end to—but that was all it was all compute, right? So, none of that even included the database resources that we were using inside of AWS. And we were able to shrink that down to a nine node Kubernetes cluster where three of those nodes were part of the control plane, and then the other six of those were part of the data plane. And then, we ended up spinning up—we were using HashiCorp Vault, and we ended up moving that outside of the cluster just for sanity purposes. But we were able to drastically decrease the footprint of an environment quite a bit, and on top of that, it also correlated to being able to decrease that spend. And so, once we started turning on everything and turning off all of the older infrastructure, it was something that we really liked. And I almost did this, and I wish I would have just taken a snapshot of those couple months within our Amazon bill and, like, posted it on a wall because it almost cut in half to the spend that we had previous to that.

Emily: And then, you mentioned some lessons. What are some top three lessons that you learned along the way?

Travis: Oh, man. So, I would say, probably the one that bit us the most was actually the telemetry, observability, being able to see what was happening within our environments, especially during a transitional time like that. Now, we did this a few years ago, and so the tools that are out now weren't necessarily as readily available as they were then. I'm not going to name companies, but the company that we were using at that point in time, we came to them and said, “We don't have this visibility, and this is hurting us. This is, kind of, a deal-breaker, and if we can't get this visibility, then we have to look elsewhere.” And they're like, “Well, it's something we've been talking about, but it's not something that we're doing right now.” And it's like, “Okay.” So, we moved on to a solution that was very much in our hands. So, we went from one to where it's like, “Well, we can't rely on a company, maybe we can just deal with it ourselves.” So, we did that, and then we realized, this is actually a lot of work, and it takes a lot of time and a lot of effort. And so, we actually stayed on that one for about a year, and then we moved off of that one, even. And where we found a middle ground to—we wanted control in certain areas, but we didn't want all of the control. And so, then we found a solution that helped us, kind of, meet a middle ground, to where we got the control we wanted, we got the flexibility using the [unintelligible] tools, primarily Prometheus, and then we were able to hand off a lot of the management of the infrastructure for the metric system and telemetry to a vendor, to where we didn't have to worry about that side of it. But we could pump over anything that we wanted, and we could aggregate the data any way we wanted, and that's exactly what we wanted to get out of that. So, that one, I think, was maybe one of the hardest ones just because we put so much work into multiple different iterations of what that eventually became. So, the one that we finally settled on to where it was, kind of, a happy medium is the one that the company is using to this day. It ended up being a much better solution, but it took us two years to figure it out.

I would say maybe the next hardest one after that is, it really comes down to just being flexible. Like, you always go in with a plan, and you always assume that that plan is going to work out, and that everything is going to be perfect. And most of the time, that doesn't end up being true. Most of the time you get to the point where you hit something, you hit a snag, or you hit some issue to where you realize that your plan is basically thrown out the window. And there was a point in time to where we, kind of, just stuck to it. We're like, “Okay, just get it to work, just get it to work, just get it to work.” And we kept trying to slam that effort moving forward until we realized that doesn't work. We're burning time. There's no way we're going to get to the point where we need to be, and we're not getting the results that we want. And so, one day I grabbed my team and we sat down and I just said, “Okay, we have this solution in place, but here's the problems with it. Of those problems, how many of them do we absolutely know how to solve right now?” And so, then we looked at the list and we talked about the ones that we knew we could solve, and it's like, okay, of the list that we don't know how to solve, there was a fair amount still leftover and looking at that list, it's like, is it going to be worthwhile to continue addressing this unknown? Or should we adapt our plan to remove that unknown piece of it, so that we can actually get back on track to what makes sense for us and for our end goals. And so, we decided it may be best to scrap that idea and go back to the drawing board. And so, we took two days to where we took an offsite. So, Nav has a corporate apartment, and we just went there and hung out there for two days, and then we whiteboarded and put post-it notes, the giant post-it pad notes, all over the walls, and then we went back to the drawing board. And this was actually around our Kubernetes management layer, what to use to help us manage Kubernetes. And so, the solution we had in place before just wasn't cutting it. And so, we went back and we literally tried everything that was available. We did a Google search, we went to any site that said, “Here's a Kubernetes management layer.” Either just a CLI to help you get the infrastructure spun up, or if it was a GUI, and it had the full management system baked in, or whatever it ended up being. So, we sat down and took that entire list, and then we took a list of the specific outcomes that we needed; we wanted to be able to do X, Y, and Z, and if any of those cannot be done with one of those solutions, then that solution is cut. And so, we seriously just took hour chunks, two-hour chunks of time, and we would divide that list up of different offerings and we started figuring out which ones would work, which ones wouldn’t until we got to the point where we literally had one left and that one ended up being the solution that we used moving forward. But being able to take that stop, and being able to readdress our plan and say, “We still want a particular outcome, but the way that we're approaching it is not working. Can we actually readdress this and change our plans in order to still get us the desired outcome?” And after we did that the one time and we got back on track a lot further than we thought, or a lot quicker than we thought we'd be able to, but after we did that the one time, then we started doing that a lot more with a lot of other issues that would arise, we would come back to it. And after we did that one time, that's when we went back to our metrics and said, “Okay, maybe we need to do the same thing with our metrics.” And that's when we shifted that the final time as well. But I think that the second one, then there really was, you have to understand that the important things out of creating a plan are the results of that plan, not necessarily how you get there. And if you're okay with changing the way that you get there, then you can actually achieve the goals of that plan much quicker.

Emily: Was there a third lesson that you learned that, sort of, stuck out?

Travis: Yeah. I'd say a third lesson is really understanding why you would want to shift over to a cloud-native infrastructure. Because at first, a lot of the reason that it started was we need to do this for cost savings. We want to be able to wrangle in our infrastructure and do all of that stuff. And it's like, okay, that's an okay reason. But at the end of the day, after I hired a team, and even after we did all the work to push everything out, were we in a net positive as far as the cost was concerned? Because there's a lot to incorporate there. And there's a lot of tools, as well, that you have to also consider, a lot of things that we ended up picking up later on that we weren't necessarily using beforehand. And so, while we were able to wrangle in and control the cost of our cloud spend, I don't know that it actually ended up being more cost-effective overall for the company. Now, that's like apples to apples, right? Let's look at our team size, and let's look at our infrastructure costs before and after. If you combine those two things together, were they less? And I don't know if that's true. But what I do think is true, is after going through all of this, we were able to move drastically faster in our pace inside of engineering. After we were done, all of the teams had their own services set up, they were able to deploy on-demand, things became very, very simplified for them. And on top of that, we even, for quite a while had a development environment that was using containerization, and it was very simple to be able to hire someone in, and just run a command, and you would have your development environment up and running. And we even had quite a few people, just the first day that they were on the job, be able to create a commit and contribute, which was a great thing. And so, if you're comparing apples to apples as far as like, what's the cost, then I don't think that that's a good reason to start addressing cloud-native infrastructure. But if you're looking at the overall cost that we had burned in engineering time trying to get development environments set up, or burned in infrastructure, trying to release a new service, or burned in many, many other ways, then we were absolutely net positive in a situation. So, releasing a service before we moved into containerization took about two weeks, and you had about four or five different people involved in order to get that service released. And it was very time consuming, and very costly. Afterwards, after we moved into containerization, it took a matter of minutes. Like, as soon as a developer wanted to release a new service, they just built out the profile in GitLab, and then they would push the code up, and it would go deploy, and everything would be up, and available, and ready to go. And so, our operating costs, I think is really what I'm coming to, is that those drastically changed. And so, at first when I was reporting about our progress and how things were going, and they kept saying, “Well, where's our cost? Where's our cost? Where’s our cost?” And so, I kept showing them, “Okay, well, this is what our infrastructure cost before and this is what it cost after.” And while there was some movement there, the thing that I started learning and started reporting back up toward the exec team later on was, “Okay, let me show you what the scenario is now as opposed to what the scenario was then.” And as soon as I was able to start painting a picture as to how much easier and faster we were able to move, they actually quit asking me. They're like, “Okay. We're good. We're sold on the fact this was successful.” And so, I think the third lesson that we learned is that it is important to understand why. And hopefully, you can figure that out before you start everything, but we didn't quite figure it out at that point in time. But we did figure it out soon enough to where we were able to make choices and adapt to that reason why to make it more beneficial for the company in the long run.

Emily: And did you feel like there was anything that was lost in translation when you were talking with the executive team and, sort of, giving updates?

Travis: Um, no, not really. I’ve had a few conversations on that and there's a lot of different things that they care about. Usually from an executive team, you want to make sure that, with what's being produced, it's not only going to be able to facilitate product movement, being able to adapt to the changes of our customers, but that we're not doing it at a pace that is, unmaintainable, which is, kind of, where we had been. And so, my conversations with them went from like, “Let's stop looking at this one particular metric that you keep asking me about, to looking at the bigger picture. How much quicker are we able to push code? How much quicker are the product owners able to adapt? How much quicker are they able to take feedback, and apply that, and put that into our product, and be able to version on top of that, and create iterations?” And on top of that, also saying, “Well yes, of course, we still have this one metric that does still matter. But that aside, I look at the overall operations that are happening now as opposed to the way that they were.” And so, for the most part, sometimes it would take a little bit of explanation, and I'm not going to lie, there are a couple times where I had to make powerpoints, and I had to, kind of, lay things out in a different way, but I think that it ended up being so well received that there was even one point where I had to present to the entire company and tell them about our migration, and what happened, and the impact that it had on our development time, and on our infrastructure costs, and everything else. Because through a migration, there are going to be pains with that migration. And after it was all said and done, the executive team wanted the entire company to understand and know why we had to go through those pains, and why it was necessary to move forward. And so, yeah, I ended up talking to the entire company and illustrating to everyone why this was such a monumental move for us. So, I don't think there was a lot lost in translation. I think it actually was very well received.

Emily: Tell me a little bit more—you were talking about how you basically were in charge of prioritizing which tools you were going to buy, what you were going to build internally. Tell me a little bit more about both what you were looking for, how you were making that decision, what some of the choices were that you made?

Travis: Yeah, for sure. So, let’s kind of like, simplify that a little bit. It comes down to—and I was able to give a talk at a couple conferences about this specific thing, but building versus buying. Why would you want to build versus why would you want to just buy? And the result that I came down to, and with a lot of help from reading a lot of information from the internet and also from some mentors, is that the most expensive resource that you have are the people that are on your team, that are working with you. Anything else above and beyond them is actually second. And so, the thing that you want to put your most expensive resource towards are the things that are going to end up evolving your company and progressing your company the fastest. And so, if there are solutions out there that you could use—or that you could build, you could like, eh, I could save a few thousand dollars and we could build this ourselves and blah, blah, blah. It's like, okay, you're looking at the purchase price of that software. But are you looking at the development time and hours that said company has put behind it, and the time and effort that you're going to end up putting behind it? Because I don't know about most companies, but everyone on my team wasn't free. There was a price behind them at the end of the day, and what they were spending that time and effort on, for me, needed to be absolutely necessary for the progression of the company. So, when I started looking at solutions, I started just deciding if we build this ourselves or if we take the time to do an open-source solution and we have to manage this ourselves, what is going to happen into the management of my team? What's going to happen to the overhead of my team? Because I can't just go hire more developers because I want to use a new open-source solution, and I need someone to maintain it. And I don't necessarily have anything against open source, but a lot of times, that's what it ended up coming down to. I think it's very valuable, and I think there are situations where it is the right way to go. Anyway, with those decisions, it really came down to if we implement this and we have to manage it, then that is time that my team is not going to be able to spend on these other projects, which those other projects are more important to me. So, then we would go and look for a vendor or a solution that could help step in and fill in the gap that we were missing. And so, given containerization, cloud-native, Kubernetes, and even Prometheus, all of those are all open source tools. But a lot of times, what we would do is use something that had that open source side to it, so that we could create a standardization, and use that standardization internally, and one that would be monitored and controlled by the community, which helped a ton. But then we have a solution on top of that, that would help bridge the gap between we don't want to manage it ourselves, or we want help managing it, or we want a solution that can step in, use this standardization, but still provide the functionality that we're looking for. And so, that is that's really where—when we were evaluating what we needed to do, then we, kind of, went through that process of can we find a way to standardize around a toolset using open source? And if so, that was great. Then we would take that and say, “Okay, now can we get help with it?” And then, that's typically the route that we would end up going.

Emily: Was there anything that you wanted to buy but couldn’t? Like, there wasn't something available?

Travis: Some of the really hard ones were actually more niche. So, I would say one of the ones that we really struggled with was on the security side. Finding solutions that worked inside of a cloud-native ecosystem as opposed to a virtual machine ecosystem, from security perspectives, were not advancing nearly as fast as some of the infrastructure tools. And so, that side of things was actually very, very complicated and hard to work with. We found some startups that were starting to address this, and we were working with them and we did purchase a solution from one of them, but we kept running into they only cover this piece, they don't cover all these other pieces, because you have intrusion detection and prevention, you also have network monitoring, and you need to have forensics running against your logs, and you also need runtime protection when your environment is up and going, and then you had the virus protection, too. And so, there wasn't anywhere that we could go to just say we need a full and complete security solution, and we want it to start now; go. So, like, being able to facilitate that part of our infrastructure was actually very complicated, and we ended up having to poke around, and use some antiquated services, and we tried to update to facilitate our needs, and some of them—I hate to say this, but some of them were even just to check a box, because within the containerized world as opposed to VM world, you're not going to get the same kind of coverage. A big one, really, is virus protection. If you look at—even if you go to Docker’s website and you read about virus protection, the only way to scan a Docker environment for viruses is to shutdown Docker, which doesn't work. You can't ever shut down Docker, because that's your entire ecosystem, so you just can't do it. But you can use immutability. You can use the fact that you created your images yourself. You can sign your images to verify that they came from a trusted source and stuff like that. And so, we ended up having to piecemeal a fair amount of that together. So, of anything, I would say that's the one thing that you can't just go out and buy right now.

Emily: I realized that we've talked a lot about pain points, but I also wanted to ask about pleasant surprises. Was there anything along the journey that went much better, was much easier than you expected?

Travis: One big one was actually the overall outcome, because we went in with one perspective of like, let's save money on infrastructure, but then realizing, through the journey, how much simpler a lot of the process became, especially for developers was a very, very pleasant surprise. And on top of that, even the developer adoption of it. I know that sometimes—and I hear a lot that it doesn't go very well for some companies, and developers don't want to learn a new technology, or whatever else, but we put a lot of time and effort upfront to educate our developers. And the adoption actually went really well for us, and that was also a very pleasant surprise. I had my defenses up, I was ready to go to war and be like, “This is happening, regardless of whether you want it or not.” And I didn't ever have to do that. As soon as we sat down and we showed them the differences in the workflow and how much quicker it was to be able to adapt and make changes to their services, as well as push new services. They were just like, “Sign me up. I'm ready to go. This is way better than anything we're doing right now.” And so, that, for me, was also another very pleasant surprise.

Emily: Can you tell me a little bit more about how this experience informs your role now at Kasten?

Travis: Yeah, so I would say there's a few things. I'd say probably the primary one is having gone through this with a company, and watching the migration, and watching all of the different struggles and the different problems you have to solve to adapt a containerized workflow has definitely influenced how I approach customers working with Kasten, but also engineering, and also the executive team as well here. And working with them, and helping them understand the things that matter to the things that didn't matter. And the things that are going to affect customers more than, maybe, they would think as well, just from my own experience and having to deal with it.

Emily: Give me some examples. What are some things that do matter versus don't matter? And where do you think there's sometimes a disconnect?

Travis: Yeah. So, you know, I'll be frank here. Kasten is definitely a Kubernetes based vendor, right? And I remember there were there a couple times—and I don't know if I want my CEO to hear this, but if he does, it's okay. There were times where I remember going to KubeCon conferences or different container-based conferences, and looking at the vendors, and just thinking, I don't know if I would ever want to do that. That never makes sense to me. But when you go up and you go talk to a vendor, you go discuss the product that they're building or whatever else, they like to show you all the flashy things, the things that really make them stand out that they're like, “Hey, we can take this process and make it crazy simpler,” or, “We can do this thing for you. We can add in this service mesh, and you're going to get all of this telemetry out of your system,” and all this craziness, or, “We can build an underlying data volume so that you can have stateful applications inside of Kubernetes. And we'll do all of this,” and it's like, every single time—not every time—but most of the time when I would talk to them, and they would give me their flashy approach and tell me, “Hey, this is all the craziness you can do.” Like, I'd go back and I talked to my team and say, “Does this make sense for us? Yeah, this is cool, but the amount of work we're going to have to put in in order to adapt that or to even use that and leverage it, what is it going to buy us? What advantage is it going to give us over what we're doing right now?” And a lot of times, it didn't end up giving a lot of advantage. It didn't make a huge difference. Now, being at one of those vendors, one of the big differences, and this was, kind of, a long-running thing with me and Niraj, our CEO, we ended up having a ton of conversation around this, but the big difference that I see with Kasten, and one thing that I continue to push here, and I told him time and time again, this is why I joined this team, is because Kasten, while they have their—we do data management, we can do backups, we can do recoveries, we have data mobility, right? The thing about Kasten is it actually lets you attack a problem the way that you want to attack it, and that's stateful applications. And a lot of times, you're going to go look in how to run stateful applications and you're going to get this big long—oh, you need a data layer. You need to be able to have your data be—to migrate across availability zones, or across regions to be able to do this. And that adds so much complexity, where at the end of the day, how often does the data infrastructure actually go bad? We have these cloud providers now, and they have spent a lot of time on making sure that their data infrastructure is pretty robust. Why aren't we just using those? Why aren't we just using those and then accounting for disasters or issues coming up around that? And that's actually the way that Kasten has approached it is, you can use your data, you can use whatever you want, and we're just here as a tool to help you facilitate that process. And so, kind of, getting back to your question of, like, what I really feel like makes a difference in this space is you have to understand what that customer is trying to do. And you have to understand how to facilitate their end goals and what they want. It's not about coming in and saying, “I can help you do all of this stuff.” And it's like, “Okay, but what does all that stuff get for me? Because really, the problem that I'm dealing with right now is, is x, y, and z.” And as a vendor, and as talking to customers, it's more about helping them. It's more about solving their problems, allowing them to focus on the tasks that are going to be more monumental for their company, instead of focusing on tasks that aren’t. Not everyone is going to be a data management company, and rightfully so. You have other things and important things to be paying attention to. So, let me come in and help you address that need without causing you a lot of pain, and a lot of hardship, to be able to just come in and use a solution and move on, but using a solution in your environment in your way to where I'm helping solve a problem, instead of helping create another problem, for what benefit?

Emily: Do you think that most companies that are on this journey are essentially trying to solve similar problems?

Travis: And which side? On the vendor side or on the consumption side?

Emily: Oh so, like Nav, the end-users. Do you think essentially any company that's moving to containers, that’s moving to Kubernetes, are they going to run into essentially the same set of problems?

Travis: You know, no, I don't think so. I think that each journey is going to be a little bit different, and it's going to cause different problems. Because if you take a company like Nav to where we had to be PCI compliant. We had different regulations that we had to abide by. And that caused the solution set that worked for us to be drastically different from the company that may not have those issues. A Kasten, for example. We're still very much, even though we have a product in the Kubernetes space, we're also still a consumer of those technologies as well. But our problems, and the things that we're addressing are monumentally different than the ones that we're addressing in Nav. And then, you also get into a lot of questions around what are the things that are important? Because sometimes your SLA is the most important thing and that will cause your solution to differ. Sometimes your SLA can waver a little bit, but you absolutely have to provide a different need for your customers. And so, while all of these tools kind of look the same—like if you're looking out in the morning, and you look at the freeway and you see all of these people that are in vehicles, and they're all traveling somewhere. Sometimes these people are moving large products. Sometimes these people are only moving themselves. But sometimes when only moving yourself, sometimes you're going to work, but sometimes you're going to play. The reason we all acquired a vehicle is because it helps facilitate that process though our need for that vehicle is drastically different. And I think that in cloud-native and Kubernetes it's the same thing. The needs are so varying and so different, but yet you can use similar tools to help facilitate them in different ways.

Emily: Do you have one or two examples of how Nav and Kasten have different needs?

Travis: Yeah, absolutely. So, I would say that one of the foremost concern that Nav is absolutely security. With the PCI regulation and everything else, protecting the identity of our customers, protecting the data for the company, it is a must. There is no if, and, or but about it, it has to happen. And the way that we ended up using Kubernetes had to facilitate that as well. So, like I had mentioned earlier that we had six nodes that we were using for the compute side of things. The reason we had six is because we had to create a logical segregation within those nodes to protect the services running on them. So, we would only allow back end services that had access to confidential information on a subset of nodes. And we wouldn't allow anything else to run there. So, you could run your front end service and a PCI compliant service on the same node, ever. But if you look at what we do at Kasten, we are running quite a few environments, and being in the Kubernetes ecosystem, and being a vendor there, we end up having to work with every single cloud vendor out there. We’re getting certified with all of them—I'm working with a few right now. But we have certifications within AWS, and Google, and Azure, and we also are working with VMware Pivotal. So, it’s across the board, and that's something that's been crazy important for Kasten is being able to have that multi-cloud experience. Being able to take data and move it from one environment to another, whether on-premise or off-premise. And so, that being one of our primary needs at Kasten. And so, we build around need, whereas Nav builds around security.

Emily: Excellent. Anything else that you'd like to add that I maybe didn't think to ask, didn’t know to ask?

Travis: Oh, that's an open-ended question. I would say one thing, if nothing else, I am very much in agreement with the fact that almost every company out there is someday going to end up hearing the words Kubernetes, just the same as they ended up learning VMware associated with virtual machines and stuff. It is. And there's a reason behind that, but the reason behind, I don't think is as important as understanding when and why it makes sense for you to start adapting and adopting those technologies because for every company, just as we've been talking about, it ends up being different, drastically different. And I think that it is very important to understand your end goal, and getting into it. Look at the overall outcome of what you're trying to achieve and use that to help drive the movement forward. Because if you look at a lot of—and the reason I say this is because if you look a lot of—I don't know if you want to call them fads or movements within technology—you look at Agile, you look at microservices, you look at a lot of other—even cloud-native. A lot of times people look at that and they're like, “Hey, look at all these good things that come out of this. And they don't typically look at what the trade-off is. Because in a microservice infrastructure, if you've got two developers, then why do you need seven different microservices? It might actually be anti—or working against what your workflow is like. And I think that even containerization is that same way. There are situations to adapt your workflow to start using those technologies, and I think sometimes there aren't. And I think that it is something that—you need to go into it and understand what the outcome is. And if you understand that outcome, then when you engage and start using those technologies, then every decision you make will be to help drive towards that outcome. And I think that that'll help you get through it a lot quicker and a lot easier, and it'll also help you just get rid of a lot of the other noise that's out there. And it'll help you, kind of, get specifically to the point of things that make sense to you and to your company so that you're able to get to that outcome and continue to drive forward and continue to help your company become successful.

Emily: All right, one last question. Actually two last questions. What is a can't-live-without engineering tool for you?

Travis: Oh, man. There's probably so many. But for me, probably one of—let me think. Is this, like, a tool that I use on my computer, or is this maybe something I use in the process? Or any of the above?

Emily: Any of the above. I mean, you could tell me Slack or something, anything that you can't imagine doing your work without.

Travis: Yeah, Slack, I think, actually helps deter me from getting work done. [laughs] But I would say for me, the one I cannot live without is a pipelining system. And for me, a lot of times has come down to GitLab. I really love the workflow in GitLab, but any pipelining system, really, is the must-have because if you can get into your process of getting code from a developer's laptop to automating how that gets into an environment, that process saves so much time and so many resources, that I don't even care which system you end up using. But just having that process, having that CI/CD system, I think is an absolute must-have.

Emily: Excellent. And then, how can people connect with you?

Travis: Yeah, I'm on some social media, I don’t do it all, but I'm definitely on LinkedIn. You can just search for me by name on there. I'm also on Twitter. My callsign there is @stmpy. It's kind of a long story, but my friends make fun of me because my legs are short, and so they used to call me Stumpy. So, it's Stumpy without the U, so just S-T-M-P-Y on Twitter. And I think that those are probably the two best ways to get a hold of me.

Emily: Excellent. Well, thank you so much for chatting.

Travis: Yeah, thank you. I really appreciate it.

Announcer: Thank you for listening to The Business of Cloud Native podcast. Keep up with the latest on the podcast at thebusinessofcloudnative.com and subscribe on iTunes, Spotify, Google Podcasts, or wherever fine podcasts are distributed. We'll see you next time.

This has been HumblePod production. Stay humble.

View Details

Some of the highlights of the show include

  • Why Adform decided to move to a cloud native architecture and Kubernetes specifically
  • Who was the driving force behind the move to Kubernetes?
  • Was the switch purely an engineering decision or did it involve people outside of engineering?
  • Positive and less positive surprises that come with switching to cloud native
  • Organizational and technical problems Edgaras has faced
  • What’s next for Adform on their cloud journey

Links

  • LinkedIn: https://www.linkedin.com/in/apsega/
  • Twitter: https://twitter.com/Apsega

Transcript
Announcer: Welcome to The Business of Cloud Native Podcast where we explore how end users talk and think about the transition to Kubernetes and cloud-native architectures.

Emily: Welcome to The Business of Cloud Native. I’m Emily Omier, your host. And I’m here today with Edgaras Apsega, lead IT systems engineer at AdForm. Edgaras, what I’d like to do is just start out with you introducing yourself.

Edgaras: I’m Edgaras. I’m working in the Adform. For anyone that doesn't know, Adform is one of the leading advertising technology companies in the world, and provides the software used by buyers and sellers to automate digital advertising. And, probably one of the most interesting parts of our solution stack is demand-side platform that has real-time bidding. And, what it means is that when that page is loading for some kind of internet users, behind the curtain, there's actually a bidding process that takes place for the placeholders to show ads. So, basically, you're doing low latency stuff. And, in Adform, I'm a lead systems engineer for the cloud services team. Our team consists of eight people, and we are providing private cloud storage, load balancing, CDN, service discovery and Kubernetes platforms for our developers that are in [00:01:36 unintelligible] production services. So, to better understand the scale that our team is working on, first of all, you can see that we are not using public cloud and we have our own private cloud that has six regions, more than 1500 physical servers, and there are more than 4000 [00:01:55 unintelligible]. And, for Kubernetes, we have seven clusters, more than 50 physical machines and around 300 constantly running [00:02:05 pods]. So, we can say that we prefer bigger clusters with bigger resources sharing pools. And you asked, how do I spend my daily work, right?

Emily: Yeah. So, when you get into the office or—right now you're not going into the office—get into your table or your [laughs] home office, what are the first couple things that you do, or…

Edgaras: Yeah, so, when I arrive at work, or, like, at these times, just get off the showers straight into work desk, [laughs] actually, I'm most productive in the mornings and evenings. So, in the mornings, when I go to my work desk, I try to do as much as I can. My sprint plan tasks, and then I scroll through the Slacks, emails, and the tickets assigned to me because we have a development team in another region. So, instantly in the mornings, we have some kinds of support tasks that we need to do.

Emily: Let's go ahead and talk about what this is all about, the business of cloud native, and tell me a little bit about why Adform decided to move to a cloud native architecture. Why did you decide to use Kubernetes, for example?

Edgaras: I'd say, actually, there were two parts. At first, we moved from traditional and, let's say, old-fashioned monitoring solutions to Prometheus, and its integration with service discovery solved lots of operational time for constantly managing and configuring monitoring and alerting for our, quite often, changing infrastructure. And the second part is the adoption of Kubernetes and all of the together coming parts like continuous integration and delivery. So, why we moved to this kind of architecture? It was because the biggest pain points for developers were to maintain actually their virtual machines. And rolling out new software releases in an old-fashioned way, took just lots of time for new software releases to reach production. So, we were looking at the new solutions that were available in the market, and Kubernetes was actually one of them. So, after successful proof of concept, we have selected it as our main application scheduler and orchestration tool.

Emily: What would you say was, like, the business value that you were hoping to get out of Kubernetes, out have the ability to release software faster, for example?

Edgaras: Yeah. So, actually, we wanted to remove the operational time from our developers so that they could spend more time coding without taking care of all of the infrastructure surrounding parts, like the application operating system management, [00:04:58 unintelligible] monitoring, alerting, logging, and so on. So, basically what, I'm saying is that the business value was for the developers to be able to ship features faster, and have a more stable platform that scales application [00:05:15 unintelligible] as well. So, in addition to that, we have a big research department, and the research department always wanted us to have a dynamic environment where they could just launch an applications around some research models, and then shut it down. So, I believe that was the business value.

Emily: Who in the organization do you think was motivating, or driving the move to Kubernetes?

Edgaras: I'd say, actually, it was more like the operation engineers, because the developers ended taking care of their environment virtual machines. They don't know much about it, but they still have to look after it, and constantly asking us for help. And we wanted to have this operational stuff only in our hands and for the developers to run only the code. So, I believe, yeah.

Emily: To what extent was the move to Kubernetes, or to cloud native in general, just purely an engineering decision? Or did it involve other people outside of engineering?

Edgaras: Well, it wasn't only the engineering decision, because we had to take it to the upper levels, just to show this new cloud native, the modern way of developing and running applications. So, the upper management level had to invest time for us to move to microservices oriented architecture and so on. So, basically, we had to show that with a little bit of time investment we can gain lots of benefits, like faster code deploys. So, we are taking the operational work from developers, and developers, when they're releasing their applications, they have full stack monitoring, logging, and they don't need to do any of the operational tasks.

Emily: How difficult was it to have this conversation? Do you feel like the upper management, did they understand the value?

Edgaras: Yeah, it was kind of hard, because nobody wants to invest time to write the code. And, as we are a software company, we always need to write new features. But, once we showed a good example, when investing not so much time, we have those kinds of benefits, then it was quite easy to change the mindset of upper management.

Emily: And, how important do you think this was for Adform?

Edgaras: I think it was very important because now what we see, we have, basically, until now we had only dozens of deployments per day. Now with Kubernetes, we have more than 500 deployments per day, which is a big number for us, and this means that we are making releases more faster.

Emily: Tell me a little bit about any surprises that you had as you were moving to Kubernetes, as you were moving to microservices. Surprises, and I'm interested in hearing both about surprises that were positive and surprises that were less positive.

Edgaras: Probably the biggest surprise for us, for our thing was just how amazing the communities. When we faced any kind of issue, most of the time there's simply a GitHub issue that’s described fixes or workarounds. You can always get an answer for questions in Slack. I remember when we had actually an issue with Kubernetes and persistent storage, and in Kubernetes Slack channel, one engineer from a company that provides storage solutions, he just provided me lots of information and several ways of tackling the problem that we were facing, actually, and that really stood me out. And, actually, we just recently started a cloud native [00:09:28 unintelligible] meetup group, with which we gather lots of folks for knowledge sharing presentations, and discussions afterwards, and it feels like the community is really strong and is eager to share their knowledge freely. So, that really amazes me about this journey.

Emily: What about some less positive surprises?

Edgaras: Yeah. So, moving to Kubernetes from virtual machines world, first of all, to change the developers mindsets about the resources utilization, I'd say. Because coming to Kubernetes world, developers need to set containers, resource limits, and often they're setting amounts similar to what they had on virtual machines with other services like monitoring, log shippers, and so on. And we see on Kubernetes, that for some applications, the resource usage is very low, but the requests of CPU is quite high, so we're still monitoring resource utilizations, and communicating with teams to lower them. Because one good example would be that while general CPU usage in whole environment is around 30 percent, we're constantly reaching fully CPU requested Kubernetes nodes, and other teams are facing deployment issues. The nodes are full. And, probably I should share that we had one interesting example that when we have migrated a service from virtual machines to Kubernetes, that service was using nine virtual machines with 16 CPUs each. And then they migrated to Kubernetes with all of the built in monitoring tools and so on. They have noticed that for the current workloads, they only needed six CPUs. So, instead of nine virtual machines with 16 CPUs, they only needed six CPUs, and so they returned just a lot of resources to the shared pool.

Emily: Wow.

Edgaras: Yeah, that's amazing. And another big pain point is always the security. So, we’re struggling a lot with the security part at the moment. And, as you may know, often security is focused on the IP address based identity, and in Kubernetes, those IPs are always changing and you can't rely on the fact that a specific IP address is tied to a particular service. So, yeah, so all the cloud native mindset needs to be changed, not only for the developers and operational engineers, but for the security engineers as well.

Emily: Where would you say you are in the cloud native transition? Are you there, have you done everything that you can, or are you somewhere on the journey?

Edgaras: I’d say, we are more than halfway through because simply [00:12:31 unintelligible] have some legacy applications that need to be rewritten so those can run in a containerized workloads. And for our critical and user-facing applications, we’re still have lots of discussions with our security team about how the infrastructure and all of those access control things should look like. So, yeah, at first, load services owners were looking at Kubernetes from a distance, and after a few successful migrations, more and more high load services are scheduled to do migrations. But in terms of the legacy applications, business still doesn't invest money, because it's not a critical application. So, I think they're going to stick for a while on that kind of phase.

Emily: Do you think that that's okay? Would you rather invest the money in—is there any disadvantages to keeping these legacy applications around?

Edgaras: I think the one point or another, they'll be completely rewritten or terminated for good. So, actually, it depends. I think, if it's not business critical, then probably it's okay. But if it is business critical, then I'd say migrate it to Kubernetes to have the self-healing infrastructure that scales just beautifully.

Emily: When you think about some of your pain points, do you think of them as technical issues? Do you think of them as, sort of, organizational issues? What are some examples of both organizational and technical problems that you've had?

Edgaras: Yeah. So, regarding the resources of the organization on how the developers are setting the resources, I think that it's kind of organizational issue. We did some Kubernetes trainings internally, and developers are always asked us to [00:14:25 unintelligible] one more time, those trainings because they're interactive. And it's a [00:14:29 hike] now. But there's always new developers coming and you still have to share your knowledge about how the resources should be implemented, how they should set the requests, or how they should set their limits and so on. Regarding the security around the Kubernetes, I think that this field is quite new, and I remember the last KubeCon in Barcelona, there were lots of buzz about the Kubernetes security and just shows that this, kind of, new field, and everybody needs information about it.

Emily: I think that you're right. It seems like both of those things are really almost skill gap issues. Do you think that there any real technical problems? So, things that the technology isn't quite there, or it's not really a problem with the way that your team members are thinking about something, or that they don't have the skills.

Edgaras: Yeah, so actually, about Kubernetes. As I mentioned, we're running Kubernetes on bare metal. And the technical stuff with Kubernetes is that it's actually first class citizen for public clouds, but when you're trying to run it on bare metal, there's some issues that you cannot expose services with, let’s say, Type LoadBalancer, you cannot have quite easily service mesh that talks not only within Kubernetes cluster, but also outside Kubernetes cluster with your virtual machines because you need to have BGP mesh, and that's your current network equipment. And there's a lot of technical issues, actually, around running Kubernetes on bare metal.

Emily: I think that that's really interesting. What are you doing to make it easier to run Kubernetes on bare metal? Or are you? Is that something that you're investing time and money into making easier?

Edgaras: Yeah. So, for running Kubernetes on bare metal, actually, we’re not using any of the automation that's provided publicly. So, we took parts of Kubernetes and automated those parts by ourselves. And we have those three data centers close to each other, connected via that fiber and we have one logical Kubernetes cluster across three data centers. And for the services to be exposed as a, let’s say, Type LoadBalancer, they do have some workarounds that will put in a custom load balancers in front of the Kubernetes cluster.

Emily: Is that something that you would hope that the community would do more of, or do you feel like you've got a pretty good handle on it at the moment?

Edgaras: Probably, I would like to see this addressed by the community more because everything that's being built for Kubernetes, it seems that it's being built for the public clouds, but not for the bare metal.

Emily: This actually, sort of, leads me to a future-oriented question. Where do you see, sort of, your next steps on the cloud journey as being?

Edgaras: Yeah, so service mesh. [laughs] everyone's talking about the service mesh. Probably, you'll have—actually we have plans to look at it and to do a proof of concept but, as I mentioned before, there are some technical issues if you want to make services talk between Kubernetes and virtual machines services between. So, looks like a journey.

Emily: What do you hope to get out of completing the journey?

Edgaras: So, service mesh, I believe, would provide this circuit-breaking, and service discovery would take us to another level. And so, I believe that when we end this journey, the scalability of our platform should improve as much as platform stability, and for the developers it would remove the operational tasks completely.

Emily: And, what do you think that that would mean in terms of the business?

Edgaras: You know, business is always looking for two things: to have stable platform for our customers, and to run infrastructure at the lowest possible costs. So, I think that the Kubernetes with container orchestration and auto-scaling solves the first problem, while the nature of shared resources in Kubernetes helps teams to achieve lower infrastructure costs.

Emily: And are you pretty happy with where you are now, with, sort of, the results that you've gotten at this stage? Would you do it over again?

Edgaras: Definitely. As I mentioned before, before Kubernetes, we had like, only tens or twenties deployments per day, now we have 500 deployments per days. And the developers are even happy with more features that they're getting. They're getting feature branch deployments, and green/blue deployments, and so on. So, for us operational engineers, there's less work to maintain everything because everything comes standardized. And for the developers, it's less operational work, and the just develop new service or feature and just push it.

Emily: Anything else that you want to add? And then I have a couple, sort of, closing questions to ask as well, but before then, is there anything else that you want to add?

Edgaras: What I'd like to add is that with Kubernetes, probably the biggest issues is with security because Kubernetes is kind of new thing. And seems like security stuff around the Kubernetes is one step behind. So, what I'd like to see is more solutions from the security perspective around Kubernetes.

Emily: So, just sort of in closing, I have a couple of fun questions. The first one is, what do you think, for you personally, and possibly organizationally, what's your can't live without engineering tool?

Edgaras: Prometheus because if you don't have monitoring, then it's like, flying the plane without any of dashboards, so you’ll crash soon.

Emily: Excellent. And then how can people connect with you?

Edgaras: LinkedIn is always open.

Emily: Are you on Twitter?

Edgaras: Yes, I am.

Emily: Fabulous. I think we can go ahead and wrap it up there, and thank you so much for chatting.

Edgaras: Cool. Thanks for having me.

Announcer: Thank you for listening to The Business of Cloud Native podcast. Keep up with the latest on the podcast at thebusinessofcloudnative.com and subscribe on iTunes, Spotify, Google Podcasts, or wherever fine podcasts are distributed. We'll see you next time.

This has been HumblePod production. Stay humble.

View Details

About Emily Omier

Emily Omier is a content strategy consultant who helps companies leverage content to build thought leadership, increase website traffic, grow their mailing list and book more demos. She has worked with CloudBees, Portworx, Plutora, Armory, and is a regular contributor for The New Stack. She graduated from the Columbia University Graduate School of Journalism and lives in Portland, Oregon.